Summary
With Qwen-Image-2.1, decoding the latent with the VAE on CUDA leaves a solid white block on the brightest highlight. Decoding the same latent on the CPU (--backend vae=cpu) is clean. Same result with the automatic tiling fallback and with explicit --vae-tiling --vae-tile-size 32x32 --vae-tile-overlap 0.5.
Environment
- Release
master-911-740c7ae, sd-master-740c7ae-bin-win-cuda12-x64.zip + cudart-sd-bin-win-cu12-x64.zip
- Windows 11, NVIDIA GeForce RTX 5060 8 GB (compute capability 12.0)
- Models:
unsloth/Qwen-Image-2.1-GGUF qwen-image-2.1-Q5_K_S.gguf, VAE unsloth/Qwen-Image-2.1-FP8 vae/qwen_image_2.1_vae_bf16.safetensors, text encoder unsloth/Qwen3-VL-8B-Instruct-GGUF Qwen3-VL-8B-Instruct-UD-Q4_K_XL.gguf
Command
sd-cli --diffusion-model qwen-image-2.1-Q5_K_S.gguf --vae qwen_image_2.1_vae_bf16.safetensors --llm Qwen3-VL-8B-Instruct-UD-Q4_K_XL.gguf -p "A cozy wine cellar with warm lantern light, wooden barrels, and a small chalkboard sign that reads \"CELLAR\"" --steps 20 --cfg-scale 1.0 --sampling-method euler -W 1024 -H 1024 --offload-to-cpu --diffusion-fa -s 42 -o out.png
Measured (pixels with every channel >= 250 inside x 470–530, y 245–290, right under the lantern flame):
| VAE decode |
white pixels |
bounding box |
| CUDA, auto tiling fallback (VAE decode ran out of memory; retrying with spatial tiling) |
497 |
(486,256)–(509,276) |
CUDA, --vae-tiling --vae-tile-size 32x32 --vae-tile-overlap 0.5 |
497 |
(486,256)–(509,276) |
CPU, --backend vae=cpu |
0 |
— |
The block is a sharp rectangle, not a blown-out highlight: the CPU decode shows the lantern base there. A cfg 6.0 run with the same seed has the same kind of block at the same spot. Looks like an overflow (fp16/bf16) in a CUDA op of the wan VAE on very bright values.
Without tiling, the CUDA decode at 1024×1024 also fails first with wan_vae segment 1/1 (graph) failed during weight preparation / vae decode compute failed before the tiling retry, even though 6.9 GB VRAM is reported free.
I can share the images or run more tests if useful.
Summary
With Qwen-Image-2.1, decoding the latent with the VAE on CUDA leaves a solid white block on the brightest highlight. Decoding the same latent on the CPU (
--backend vae=cpu) is clean. Same result with the automatic tiling fallback and with explicit--vae-tiling --vae-tile-size 32x32 --vae-tile-overlap 0.5.Environment
master-911-740c7ae,sd-master-740c7ae-bin-win-cuda12-x64.zip+cudart-sd-bin-win-cu12-x64.zipunsloth/Qwen-Image-2.1-GGUFqwen-image-2.1-Q5_K_S.gguf, VAEunsloth/Qwen-Image-2.1-FP8vae/qwen_image_2.1_vae_bf16.safetensors, text encoderunsloth/Qwen3-VL-8B-Instruct-GGUFQwen3-VL-8B-Instruct-UD-Q4_K_XL.ggufCommand
Measured (pixels with every channel >= 250 inside x 470–530, y 245–290, right under the lantern flame):
--vae-tiling --vae-tile-size 32x32 --vae-tile-overlap 0.5--backend vae=cpuThe block is a sharp rectangle, not a blown-out highlight: the CPU decode shows the lantern base there. A cfg 6.0 run with the same seed has the same kind of block at the same spot. Looks like an overflow (fp16/bf16) in a CUDA op of the wan VAE on very bright values.
Without tiling, the CUDA decode at 1024×1024 also fails first with
wan_vae segment 1/1 (graph) failed during weight preparation/vae decode compute failedbefore the tiling retry, even though 6.9 GB VRAM is reported free.I can share the images or run more tests if useful.