Instructions to use Felldude/QWEN_2.1_HDR_VAE with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use Felldude/QWEN_2.1_HDR_VAE with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Felldude/QWEN_2.1_HDR_VAE", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Qwen Image 2.1 HDR VAE
Fine-tuned Qwen Image 2.1 autoencoder for RGBA / HDR-oriented reconstruction.
The goal is stronger dynamic range, color volume, and detail energy for HDR-style workflows—not perfect fidelity to the source frame.
Eval
Relative to ground truth and compared with base Qwen 2.1 VAE (VAE A):
| Metric | VAE A (Qwen 2.1 VAE) | VAE B (HDR) | vs GT |
|---|---|---|---|
| Lab color volume ratio | ~0.81× | ~1.34× | +34% color volume |
| Edge density ratio | ~0.95× | ~1.50× | +50% edge density |
| Gradient strength ratio | ~0.97× | ~1.16× | +16% gradient energy |
| Sharpness ratio | ~0.96× | ~1.99× | ~2× sharpness |
| Saturation ratio | ~0.99× | ~1.20× | +20% saturation |
| SSIM | 0.958 | 0.944 | Slightly lower (expected) |
| PSNR | 38.1 dB | 32.0 dB | Lower (expected) |
VAE B deliberately overshoots GT on Lab color volume, edge density, and gradient energy. VAE A stays closer to GT on SSIM/PSNR but compresses color volume and softens detail. That tradeoff is the point of this checkpoint.
Technical notes
- Base: Qwen Image 2.1 VAE (
z_dim=64, 16× spatial compression) - Input / output: RGBA, normalized to
[-1, 1] - Layout:
[B, 4, 1, H, W](single-frame temporal dim) - Uses the VAE’s own
latents_mean/latents_std(no FLUX/SD0.18215scaling) - Trained / evaluated around 512–1024 resolution; spatial size should be divisible by 16
Intended use
- HDR-oriented encode/decode in Qwen Image pipelines
- Image editing / generation stacks that benefit from punchier latents
- Research on RGBA + high dynamic range autoencoders
Not intended for: archival reconstruction, metrics-maximizing ablations, or any setting where matching the input pixel-for-pixel matters more than look.
Limitations
- Will diverge from the source on sharpness, saturation, and color volume by design
- Alpha is supported; behavior on unusual alpha mattes is not heavily tuned
- Very large resolutions are memory-heavy (16× latent downscale)
License
See LICENSE (Qwen Research terms).
- Downloads last month
- 1,607
Model tree for Felldude/QWEN_2.1_HDR_VAE
Base model
Qwen/Qwen-Image-2.1