Git commit
88411ef (master-908). Also reproduced on 137f740 (master-883, the first Qwen-Image-2.1 commit): same seed and settings give pixel-identical output.
Operating System & Version
Ubuntu 26.04.1 LTS, kernel 7.0.0-31. NVIDIA GeForce RTX 5080 (16 GB), driver 595.91.07, CUDA 13.2.
GGML backends
CUDA and Vulkan (both affected).
Command-line arguments used
sd-cli --diffusion-model qwen_image_2.1-Q8_0.gguf --vae qwen_image_2.1_vae_bf16.safetensors --llm Qwen3VL-8B-Instruct-Q8_0.gguf --sampling-method euler --cfg-scale 5 --steps 40 -W 2048 -H 2048 --diffusion-fa --offload-to-cpu --sage-attn -s 757604577 -p "a photograph of an old stone fishing harbor at golden hour, wooden boats moored at the quay, calm water with reflections, weathered houses, sharp focus, 50mm lens, natural light"
The diffusion model is qwen_image_2.1-Q8_0.gguf from leejet/Qwen-Image-2.1-GGUF. The VAE is qwen_image_2.1_vae_bf16.safetensors.
Steps to reproduce
- Run the command above at 2048x2048, one of the model's native sizes.
- View the output at 100%.
What you expected to happen
A clean image at the model's native 2K resolution. The same prompt at 1024x1024 comes out clean.
What actually happened
Everything at 2048x2048 and 2752x1536 carries a regular fine grid, a crosshatch of horizontal and vertical lines with an 8 px period and a strong 4 px harmonic. It covers the whole image and is most visible on smooth or textured areas such as water, stone and sails. Scaled down, the images look fine; at 100% the grid is obvious, and it also shows up in painterly prompts. At 1024x1024 it is essentially absent.
Images (1024 vs 2048 crop, and a grid of 2048 variants): https://gist.github.com/bhoult/0b9535a30fd7b2ec8c4f664935cd0004


To measure it rather than judge by eye, I used the gridscore.py in the gist. It takes the power spectrum of the high-passed luminance and divides the power at frequency N/P by the median of nearby frequencies, taking the larger of the horizontal and vertical values. Around 1-10 means no grid.
| Image (same prompt; 2048 runs use the same seed) |
P=4 px |
P=8 px |
| 1024x1024, CUDA, Q8_0, cfg 5 |
1.2 |
8.9 |
| 2048x2048, CUDA, Q8_0, cfg 5, FA + sage (baseline) |
544 |
266 |
same, VAE decoded on CPU (--backend vae=cpu, no tiling) |
544 |
273 |
same, without --sage-attn |
543 |
230 |
| same, master-883 (137f740), no sage |
543 |
230 (pixel-identical to the line above) |
| same, Q6_K diffusion model |
809 |
273 |
| same, cfg 1 / 3 / 6 |
274 / 728 / 416 |
149 / 317 / 166 |
| same, Vulkan backend |
695 |
217 |
| 2752x1536 (16:9), different prompt |
295 |
268 |
What this rules out:
- VAE tiling. At 2048 the GPU VAE decode runs out of memory and falls back to tiling, but decoding the same latent untiled on the CPU gives the same grid.
- SageAttention.
- The master-883 to master-908 changes. That range includes the circular RoPE refactor and the Wan VAE 2D-convolution change.
- Quantization. Q8_0 and Q6_K both show it.
- CFG scale.
- The GPU backend. CUDA and Vulkan both show it.
Not tested:
- Running without
--diffusion-fa. At 2048 that needs about 34 GB of VRAM, more than this card has.
- Comparing against the reference diffusers pipeline at 2048. So I can't rule out that the upstream model itself does this. But 2048x2048 is listed as a native resolution on the Qwen/Qwen-Image-2.1 model card, so I'd expect a clean result there.
Since the grid is backend- and quantization-independent and appears only at the 2K sizes, my guess is shared code that depends on resolution: positional embedding or RoPE scaling for large token grids, the automatically chosen resolution-dependent flow shift, or patchify/unpatchify. The 8 px period with a 4 px harmonic may point at the patch or latent layout. This is only a guess.
Logs / error messages / stack trace
No errors from the diffusion step. At 2048 the log contains VAE decode ran out of memory; retrying with spatial tiling, which is expected on a 16 GB card. The CPU-decode test above shows the tiling isn't the cause.
Additional context / environment details
- Text encoder: Qwen3-VL-8B-Instruct Q8_0 GGUF.
--offload-to-cpu is set in every 2048 run. It only moves weights, so I don't expect it to matter.
- 3072x3072 at 20 steps also runs, but the tiled VAE decode fails there: it wants about 23 GB even when tiled.
Git commit
88411ef (master-908). Also reproduced on 137f740 (master-883, the first Qwen-Image-2.1 commit): same seed and settings give pixel-identical output.
Operating System & Version
Ubuntu 26.04.1 LTS, kernel 7.0.0-31. NVIDIA GeForce RTX 5080 (16 GB), driver 595.91.07, CUDA 13.2.
GGML backends
CUDA and Vulkan (both affected).
Command-line arguments used
The diffusion model is
qwen_image_2.1-Q8_0.gguffrom leejet/Qwen-Image-2.1-GGUF. The VAE isqwen_image_2.1_vae_bf16.safetensors.Steps to reproduce
What you expected to happen
A clean image at the model's native 2K resolution. The same prompt at 1024x1024 comes out clean.
What actually happened
Everything at 2048x2048 and 2752x1536 carries a regular fine grid, a crosshatch of horizontal and vertical lines with an 8 px period and a strong 4 px harmonic. It covers the whole image and is most visible on smooth or textured areas such as water, stone and sails. Scaled down, the images look fine; at 100% the grid is obvious, and it also shows up in painterly prompts. At 1024x1024 it is essentially absent.
Images (1024 vs 2048 crop, and a grid of 2048 variants): https://gist.github.com/bhoult/0b9535a30fd7b2ec8c4f664935cd0004
To measure it rather than judge by eye, I used the
gridscore.pyin the gist. It takes the power spectrum of the high-passed luminance and divides the power at frequency N/P by the median of nearby frequencies, taking the larger of the horizontal and vertical values. Around 1-10 means no grid.--backend vae=cpu, no tiling)--sage-attnWhat this rules out:
Not tested:
--diffusion-fa. At 2048 that needs about 34 GB of VRAM, more than this card has.Since the grid is backend- and quantization-independent and appears only at the 2K sizes, my guess is shared code that depends on resolution: positional embedding or RoPE scaling for large token grids, the automatically chosen resolution-dependent flow shift, or patchify/unpatchify. The 8 px period with a 4 px harmonic may point at the patch or latent layout. This is only a guess.
Logs / error messages / stack trace
No errors from the diffusion step. At 2048 the log contains
VAE decode ran out of memory; retrying with spatial tiling, which is expected on a 16 GB card. The CPU-decode test above shows the tiling isn't the cause.Additional context / environment details
--offload-to-cpuis set in every 2048 run. It only moves weights, so I don't expect it to matter.