Qwen Image 2.1 HDR VAE

Fine-tuned Qwen Image 2.1 autoencoder for RGBA / HDR-oriented reconstruction.

The goal is stronger dynamic range, color volume, and detail energy for HDR-style workflows—not perfect fidelity to the source frame.

Eval

Relative to ground truth and compared with base Qwen 2.1 VAE (VAE A):

Metric VAE A (Qwen 2.1 VAE) VAE B (HDR) vs GT
Lab color volume ratio ~0.81× ~1.34× +34% color volume
Edge density ratio ~0.95× ~1.50× +50% edge density
Gradient strength ratio ~0.97× ~1.16× +16% gradient energy
Sharpness ratio ~0.96× ~1.99× ~2× sharpness
Saturation ratio ~0.99× ~1.20× +20% saturation
SSIM 0.958 0.944 Slightly lower (expected)
PSNR 38.1 dB 32.0 dB Lower (expected)

VAE B deliberately overshoots GT on Lab color volume, edge density, and gradient energy. VAE A stays closer to GT on SSIM/PSNR but compresses color volume and softens detail. That tradeoff is the point of this checkpoint.

Technical notes

  • Base: Qwen Image 2.1 VAE (z_dim=64, 16× spatial compression)
  • Input / output: RGBA, normalized to [-1, 1]
  • Layout: [B, 4, 1, H, W] (single-frame temporal dim)
  • Uses the VAE’s own latents_mean / latents_std (no FLUX/SD 0.18215 scaling)
  • Trained / evaluated around 512–1024 resolution; spatial size should be divisible by 16

Intended use

  • HDR-oriented encode/decode in Qwen Image pipelines
  • Image editing / generation stacks that benefit from punchier latents
  • Research on RGBA + high dynamic range autoencoders

Not intended for: archival reconstruction, metrics-maximizing ablations, or any setting where matching the input pixel-for-pixel matters more than look.

Limitations

  • Will diverge from the source on sharpness, saturation, and color volume by design
  • Alpha is supported; behavior on unusual alpha mattes is not heavily tuned
  • Very large resolutions are memory-heavy (16× latent downscale)

License

See LICENSE (Qwen Research terms).

Downloads last month
1,607
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Felldude/QWEN_2.1_HDR_VAE

Finetuned
(55)
this model