Skip to content

[Bug] Qwen-Image-2.1: split-MLP LoRAs apply partially (326/454), tiling defaults fail, white-block notes #2051

Description

@vhanla

Git commit

b167b94

Operating System & Version

CachyOS -> Kernel: Linux 7.2.7-1-cachyos

GGML backends

CUDA

Command-line arguments used

sd-cli --diffusion-model qwen_image_2.1-Q2_K.gguf --vae qwen_image_2.1_vae_bf16.safetensors \ --llm Qwen3-VL-8B-Q2_K.gguf --lora-model-dir ./lora \ -p "a lovely catlora:Qwen-Image-2.1-viggle-turbo-4step-lora-r64:1.0" \ --cfg-scale 1.2 --steps 4 --sampling-method euler \ --offload-to-cpu --diffusion-fa -o out.png

Steps to reproduce

Setup: Qwen-Image-2.1 GGUF with fused img_mlp.gate_up (Q2_K and Q4_K tested),
qwen_image_2.1_vae_bf16.safetensors, Qwen3-VL 8B GGUF, 4 GB VRAM (GTX 1650,
sm_75), base commit b167b94.

Bug 1 : split LoRA silently partial:

  1. Run the command below (paths shortened).
  2. Check the log for the LoRA stat line.

sd-cli --diffusion-model qwen_image_2.1-Q2_K.gguf --vae qwen_image_2.1_vae_bf16.safetensors
--llm Qwen3-VL-8B-Q2_K.gguf --lora-model-dir ./lora
-p "a lovely catlora:Qwen-Image-2.1-viggle-turbo-4step-lora-r64:1.0"
--cfg-scale 1.2 --steps 4 --sampling-method euler
--offload-to-cpu --diffusion-fa -o out.png

Bug 2 : tiling defaults fail:

  1. Same base command + --vae-tiling (no explicit tile sizes).
  2. Observe the VAE decode outcome.

What you expected to happen

  1. All 454 LoRA tensors apply (the file ships split img_mlp.proj/gate_layer
    pairs for a fused gate_up target; same pattern the SD1 qkv -> in_proj rename
    already handles in preprocess_lora_tensors).
  2. --vae-tiling with defaults tiles the decode (e.g. half-latent like the
    reactive retry does) instead of failing.

What actually happened

  1. Only 326/454 apply; the 128 split MLP tensors are skipped as "unused"
    with no error, degrading few-step quality.
  2. The flag produces a single full-frame tile (32x32, 1 tile at 512px),
    OOMs exactly like untiled decode, and the retry refuses because tiling
    is already enabled => rc=1, no image.

Logs / error messages / stack trace

Bug 1:
[WARN] lora.hpp:967 - unused lora tensor |lora.model...transformer_blocks.9.img_mlp.proj.weight.lora_up|
... (128 lines, all img_mlp proj/gate_layer across 32 blocks)
[WARN] lora.hpp:979 - Only (326 / 454) LoRA tensors have been applied

Bug 2:
[VERBOSE] vae.hpp:285 - VAE Tile size: 32x32
[VERBOSE] tiling.cpp:201 - num tiles : 1, 1
[WARN] model_manager.cpp:1826 - ... need 3445.32 MB ... available 2679.19 MB ...
[ERROR] vae.hpp:136 - vae decode compute failed while processing a tile
[ERROR] image.cpp:607 - decode_first_stage failed for latent 1

Additional context / environment details

Summary

Two related findings on Qwen-Image-2.1, both reproduced on sm_75 / 4 GB VRAM and verified before/after with local patches, my patches are below in the referenced branch at my fork:

  1. Split img_mlp LoRA pairs are silently dropped against fused gate_up GGUFs (326/454).
  2. --vae-tiling with default sizes fails outright (rc=1) where the reactive retry succeeds.

Environment

  • Linux x86_64, GTX 1650 4 GB (Turing, sm_75), CUDA backend, static build (SD_CUDA=ON, ARCH=75)
  • Base commit b167b94; findings 1–2 verified before/after with local patches (branch linked below)
  • Models: qwen_image_2.1 GGUF (fused img_mlp.gate_up; Q2_K and Q4_K), qwen_image_2.1_vae_bf16.safetensors, Qwen3-VL 8B GGUF, community LoRAs (4-step turbo r64 + 13 style LoRAs)

1. Split-MLP LoRA silently partial

Command (paths shortened):

sd-cli --diffusion-model qwen_image_2.1-Q2_K.gguf --vae qwen_image_2.1_vae_bf16.safetensors \
  --llm Qwen3-VL-8B-Q2_K.gguf --lora-model-dir ./lora \
  -p "a lovely cat<lora:Qwen-Image-2.1-viggle-turbo-4step-lora-r64:1.0>" \
  --cfg-scale 1.2 --steps 4 --sampling-method euler \
  --offload-to-cpu --diffusion-fa -o out.png

Measured:

Build Applied Unused
pristine b167b94 326 / 454 128 × img_mlp.proj/gate_layer (all 32 blocks)
+ local fusion patch 454 / 454 0

13 style LoRA files (from CivitAI for Qwen Image 2.1) show the same pattern (384/384 or 448/448 after resolving; one SD3-style file correctly rejected at 0/1424 + incompatibles). Same-seed output is visibly sharper with full application; sampling cost +2–5%.

2. --vae-tiling defaults fail (rc=1)

Same base command + --vae-tiling:

[VERBOSE] vae.hpp:285 - VAE Tile size: 32x32
[VERBOSE] tiling.cpp:201 - num tiles : 1, 1
[WARN] model_manager.cpp:1826 - ... need 3445.32 MB ... available 2679.19 MB ...
[ERROR] vae.hpp:136 - vae decode compute failed while processing a tile
[ERROR] image.cpp:607 - decode_first_stage failed for latent 1

Defaults degrade to one full-frame tile => same OOM as untiled, and the retry refuses (already enabled). Without the flag, the reactive retry (16x16, 9 tiles, ~38s) succeeds. Defaulting empty sizes to the retry's half-latent tiling fixes it (rc=0, same decode time).

References

Branch with both local fixes (verified before/after):
https://github.com/vhanla/stable-diffusion.cpp/tree/qwen21-lowvram-lora-fused-mlp-vae-tiling-fix
Viggle Turbo 4, 5, 6 step LoRA
https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions