Git commit
b167b94
Operating System & Version
CachyOS -> Kernel: Linux 7.2.7-1-cachyos
GGML backends
CUDA
Command-line arguments used
sd-cli --diffusion-model qwen_image_2.1-Q2_K.gguf --vae qwen_image_2.1_vae_bf16.safetensors \ --llm Qwen3-VL-8B-Q2_K.gguf --lora-model-dir ./lora \ -p "a lovely catlora:Qwen-Image-2.1-viggle-turbo-4step-lora-r64:1.0" \ --cfg-scale 1.2 --steps 4 --sampling-method euler \ --offload-to-cpu --diffusion-fa -o out.png
Steps to reproduce
Setup: Qwen-Image-2.1 GGUF with fused img_mlp.gate_up (Q2_K and Q4_K tested),
qwen_image_2.1_vae_bf16.safetensors, Qwen3-VL 8B GGUF, 4 GB VRAM (GTX 1650,
sm_75), base commit b167b94.
Bug 1 : split LoRA silently partial:
- Run the command below (paths shortened).
- Check the log for the LoRA stat line.
sd-cli --diffusion-model qwen_image_2.1-Q2_K.gguf --vae qwen_image_2.1_vae_bf16.safetensors
--llm Qwen3-VL-8B-Q2_K.gguf --lora-model-dir ./lora
-p "a lovely catlora:Qwen-Image-2.1-viggle-turbo-4step-lora-r64:1.0"
--cfg-scale 1.2 --steps 4 --sampling-method euler
--offload-to-cpu --diffusion-fa -o out.png
Bug 2 : tiling defaults fail:
- Same base command +
--vae-tiling (no explicit tile sizes).
- Observe the VAE decode outcome.
What you expected to happen
- All 454 LoRA tensors apply (the file ships split
img_mlp.proj/gate_layer
pairs for a fused gate_up target; same pattern the SD1 qkv -> in_proj rename
already handles in preprocess_lora_tensors).
--vae-tiling with defaults tiles the decode (e.g. half-latent like the
reactive retry does) instead of failing.
What actually happened
- Only 326/454 apply; the 128 split MLP tensors are skipped as "unused"
with no error, degrading few-step quality.
- The flag produces a single full-frame tile (32x32, 1 tile at 512px),
OOMs exactly like untiled decode, and the retry refuses because tiling
is already enabled => rc=1, no image.
Logs / error messages / stack trace
Bug 1:
[WARN] lora.hpp:967 - unused lora tensor |lora.model...transformer_blocks.9.img_mlp.proj.weight.lora_up|
... (128 lines, all img_mlp proj/gate_layer across 32 blocks)
[WARN] lora.hpp:979 - Only (326 / 454) LoRA tensors have been applied
Bug 2:
[VERBOSE] vae.hpp:285 - VAE Tile size: 32x32
[VERBOSE] tiling.cpp:201 - num tiles : 1, 1
[WARN] model_manager.cpp:1826 - ... need 3445.32 MB ... available 2679.19 MB ...
[ERROR] vae.hpp:136 - vae decode compute failed while processing a tile
[ERROR] image.cpp:607 - decode_first_stage failed for latent 1
Additional context / environment details
Summary
Two related findings on Qwen-Image-2.1, both reproduced on sm_75 / 4 GB VRAM and verified before/after with local patches, my patches are below in the referenced branch at my fork:
- Split
img_mlp LoRA pairs are silently dropped against fused gate_up GGUFs (326/454).
--vae-tiling with default sizes fails outright (rc=1) where the reactive retry succeeds.
Environment
- Linux x86_64, GTX 1650 4 GB (Turing, sm_75), CUDA backend, static build (
SD_CUDA=ON, ARCH=75)
- Base commit
b167b94; findings 1–2 verified before/after with local patches (branch linked below)
- Models:
qwen_image_2.1 GGUF (fused img_mlp.gate_up; Q2_K and Q4_K), qwen_image_2.1_vae_bf16.safetensors, Qwen3-VL 8B GGUF, community LoRAs (4-step turbo r64 + 13 style LoRAs)
1. Split-MLP LoRA silently partial
Command (paths shortened):
sd-cli --diffusion-model qwen_image_2.1-Q2_K.gguf --vae qwen_image_2.1_vae_bf16.safetensors \
--llm Qwen3-VL-8B-Q2_K.gguf --lora-model-dir ./lora \
-p "a lovely cat<lora:Qwen-Image-2.1-viggle-turbo-4step-lora-r64:1.0>" \
--cfg-scale 1.2 --steps 4 --sampling-method euler \
--offload-to-cpu --diffusion-fa -o out.png
Measured:
| Build |
Applied |
Unused |
pristine b167b94 |
326 / 454 |
128 × img_mlp.proj/gate_layer (all 32 blocks) |
| + local fusion patch |
454 / 454 |
0 |
13 style LoRA files (from CivitAI for Qwen Image 2.1) show the same pattern (384/384 or 448/448 after resolving; one SD3-style file correctly rejected at 0/1424 + incompatibles). Same-seed output is visibly sharper with full application; sampling cost +2–5%.
2. --vae-tiling defaults fail (rc=1)
Same base command + --vae-tiling:
[VERBOSE] vae.hpp:285 - VAE Tile size: 32x32
[VERBOSE] tiling.cpp:201 - num tiles : 1, 1
[WARN] model_manager.cpp:1826 - ... need 3445.32 MB ... available 2679.19 MB ...
[ERROR] vae.hpp:136 - vae decode compute failed while processing a tile
[ERROR] image.cpp:607 - decode_first_stage failed for latent 1
Defaults degrade to one full-frame tile => same OOM as untiled, and the retry refuses (already enabled). Without the flag, the reactive retry (16x16, 9 tiles, ~38s) succeeds. Defaulting empty sizes to the retry's half-latent tiling fixes it (rc=0, same decode time).
References
Branch with both local fixes (verified before/after):
https://github.com/vhanla/stable-diffusion.cpp/tree/qwen21-lowvram-lora-fused-mlp-vae-tiling-fix
Viggle Turbo 4, 5, 6 step LoRA
https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo
Git commit
b167b94
Operating System & Version
CachyOS -> Kernel: Linux 7.2.7-1-cachyos
GGML backends
CUDA
Command-line arguments used
sd-cli --diffusion-model qwen_image_2.1-Q2_K.gguf --vae qwen_image_2.1_vae_bf16.safetensors \ --llm Qwen3-VL-8B-Q2_K.gguf --lora-model-dir ./lora \ -p "a lovely catlora:Qwen-Image-2.1-viggle-turbo-4step-lora-r64:1.0" \ --cfg-scale 1.2 --steps 4 --sampling-method euler \ --offload-to-cpu --diffusion-fa -o out.png
Steps to reproduce
Setup: Qwen-Image-2.1 GGUF with fused
img_mlp.gate_up(Q2_K and Q4_K tested),qwen_image_2.1_vae_bf16.safetensors, Qwen3-VL 8B GGUF, 4 GB VRAM (GTX 1650,sm_75), base commit
b167b94.Bug 1 : split LoRA silently partial:
sd-cli --diffusion-model qwen_image_2.1-Q2_K.gguf --vae qwen_image_2.1_vae_bf16.safetensors
--llm Qwen3-VL-8B-Q2_K.gguf --lora-model-dir ./lora
-p "a lovely catlora:Qwen-Image-2.1-viggle-turbo-4step-lora-r64:1.0"
--cfg-scale 1.2 --steps 4 --sampling-method euler
--offload-to-cpu --diffusion-fa -o out.png
Bug 2 : tiling defaults fail:
--vae-tiling(no explicit tile sizes).What you expected to happen
img_mlp.proj/gate_layerpairs for a fused
gate_uptarget; same pattern the SD1 qkv -> in_proj renamealready handles in
preprocess_lora_tensors).--vae-tilingwith defaults tiles the decode (e.g. half-latent like thereactive retry does) instead of failing.
What actually happened
with no error, degrading few-step quality.
OOMs exactly like untiled decode, and the retry refuses because tiling
is already enabled => rc=1, no image.
Logs / error messages / stack trace
Bug 1:
[WARN] lora.hpp:967 - unused lora tensor |lora.model...transformer_blocks.9.img_mlp.proj.weight.lora_up|
... (128 lines, all img_mlp proj/gate_layer across 32 blocks)
[WARN] lora.hpp:979 - Only (326 / 454) LoRA tensors have been applied
Bug 2:
[VERBOSE] vae.hpp:285 - VAE Tile size: 32x32
[VERBOSE] tiling.cpp:201 - num tiles : 1, 1
[WARN] model_manager.cpp:1826 - ... need 3445.32 MB ... available 2679.19 MB ...
[ERROR] vae.hpp:136 - vae decode compute failed while processing a tile
[ERROR] image.cpp:607 - decode_first_stage failed for latent 1
Additional context / environment details
Summary
Two related findings on Qwen-Image-2.1, both reproduced on sm_75 / 4 GB VRAM and verified before/after with local patches, my patches are below in the referenced branch at my fork:
img_mlpLoRA pairs are silently dropped against fusedgate_upGGUFs (326/454).--vae-tilingwith default sizes fails outright (rc=1) where the reactive retry succeeds.Environment
SD_CUDA=ON,ARCH=75)b167b94; findings 1–2 verified before/after with local patches (branch linked below)qwen_image_2.1GGUF (fusedimg_mlp.gate_up; Q2_K and Q4_K),qwen_image_2.1_vae_bf16.safetensors, Qwen3-VL 8B GGUF, community LoRAs (4-step turbo r64 + 13 style LoRAs)1. Split-MLP LoRA silently partial
Command (paths shortened):
sd-cli --diffusion-model qwen_image_2.1-Q2_K.gguf --vae qwen_image_2.1_vae_bf16.safetensors \ --llm Qwen3-VL-8B-Q2_K.gguf --lora-model-dir ./lora \ -p "a lovely cat<lora:Qwen-Image-2.1-viggle-turbo-4step-lora-r64:1.0>" \ --cfg-scale 1.2 --steps 4 --sampling-method euler \ --offload-to-cpu --diffusion-fa -o out.pngMeasured:
b167b94img_mlp.proj/gate_layer(all 32 blocks)13 style LoRA files (from CivitAI for Qwen Image 2.1) show the same pattern (384/384 or 448/448 after resolving; one SD3-style file correctly rejected at 0/1424 + incompatibles). Same-seed output is visibly sharper with full application; sampling cost +2–5%.
2.
--vae-tilingdefaults fail (rc=1)Same base command +
--vae-tiling:Defaults degrade to one full-frame tile => same OOM as untiled, and the retry refuses (already enabled). Without the flag, the reactive retry (16x16, 9 tiles, ~38s) succeeds. Defaulting empty sizes to the retry's half-latent tiling fixes it (rc=0, same decode time).
References
Branch with both local fixes (verified before/after):
https://github.com/vhanla/stable-diffusion.cpp/tree/qwen21-lowvram-lora-fused-mlp-vae-tiling-fix
Viggle Turbo 4, 5, 6 step LoRA
https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo