Skip to content

[Bug] Qwen-Image-2.1: periodic 8px/4px grid artifacts at native 2K resolutions (CUDA and Vulkan, Q8_0 and Q6_K) #2041

Description

@bhoult

Git commit

88411ef (master-908). Also reproduced on 137f740 (master-883, the first Qwen-Image-2.1 commit): same seed and settings give pixel-identical output.

Operating System & Version

Ubuntu 26.04.1 LTS, kernel 7.0.0-31. NVIDIA GeForce RTX 5080 (16 GB), driver 595.91.07, CUDA 13.2.

GGML backends

CUDA and Vulkan (both affected).

Command-line arguments used

sd-cli --diffusion-model qwen_image_2.1-Q8_0.gguf --vae qwen_image_2.1_vae_bf16.safetensors --llm Qwen3VL-8B-Instruct-Q8_0.gguf --sampling-method euler --cfg-scale 5 --steps 40 -W 2048 -H 2048 --diffusion-fa --offload-to-cpu --sage-attn -s 757604577 -p "a photograph of an old stone fishing harbor at golden hour, wooden boats moored at the quay, calm water with reflections, weathered houses, sharp focus, 50mm lens, natural light"

The diffusion model is qwen_image_2.1-Q8_0.gguf from leejet/Qwen-Image-2.1-GGUF. The VAE is qwen_image_2.1_vae_bf16.safetensors.

Steps to reproduce

  1. Run the command above at 2048x2048, one of the model's native sizes.
  2. View the output at 100%.

What you expected to happen

A clean image at the model's native 2K resolution. The same prompt at 1024x1024 comes out clean.

What actually happened

Everything at 2048x2048 and 2752x1536 carries a regular fine grid, a crosshatch of horizontal and vertical lines with an 8 px period and a strong 4 px harmonic. It covers the whole image and is most visible on smooth or textured areas such as water, stone and sails. Scaled down, the images look fine; at 100% the grid is obvious, and it also shows up in painterly prompts. At 1024x1024 it is essentially absent.

Images (1024 vs 2048 crop, and a grid of 2048 variants): https://gist.github.com/bhoult/0b9535a30fd7b2ec8c4f664935cd0004

1024 vs 2048

2048 variants

To measure it rather than judge by eye, I used the gridscore.py in the gist. It takes the power spectrum of the high-passed luminance and divides the power at frequency N/P by the median of nearby frequencies, taking the larger of the horizontal and vertical values. Around 1-10 means no grid.

Image (same prompt; 2048 runs use the same seed) P=4 px P=8 px
1024x1024, CUDA, Q8_0, cfg 5 1.2 8.9
2048x2048, CUDA, Q8_0, cfg 5, FA + sage (baseline) 544 266
same, VAE decoded on CPU (--backend vae=cpu, no tiling) 544 273
same, without --sage-attn 543 230
same, master-883 (137f740), no sage 543 230 (pixel-identical to the line above)
same, Q6_K diffusion model 809 273
same, cfg 1 / 3 / 6 274 / 728 / 416 149 / 317 / 166
same, Vulkan backend 695 217
2752x1536 (16:9), different prompt 295 268

What this rules out:

  • VAE tiling. At 2048 the GPU VAE decode runs out of memory and falls back to tiling, but decoding the same latent untiled on the CPU gives the same grid.
  • SageAttention.
  • The master-883 to master-908 changes. That range includes the circular RoPE refactor and the Wan VAE 2D-convolution change.
  • Quantization. Q8_0 and Q6_K both show it.
  • CFG scale.
  • The GPU backend. CUDA and Vulkan both show it.

Not tested:

  • Running without --diffusion-fa. At 2048 that needs about 34 GB of VRAM, more than this card has.
  • Comparing against the reference diffusers pipeline at 2048. So I can't rule out that the upstream model itself does this. But 2048x2048 is listed as a native resolution on the Qwen/Qwen-Image-2.1 model card, so I'd expect a clean result there.

Since the grid is backend- and quantization-independent and appears only at the 2K sizes, my guess is shared code that depends on resolution: positional embedding or RoPE scaling for large token grids, the automatically chosen resolution-dependent flow shift, or patchify/unpatchify. The 8 px period with a 4 px harmonic may point at the patch or latent layout. This is only a guess.

Logs / error messages / stack trace

No errors from the diffusion step. At 2048 the log contains VAE decode ran out of memory; retrying with spatial tiling, which is expected on a 16 GB card. The CPU-decode test above shows the tiling isn't the cause.

Additional context / environment details

  • Text encoder: Qwen3-VL-8B-Instruct Q8_0 GGUF.
  • --offload-to-cpu is set in every 2048 run. It only moves weights, so I don't expect it to matter.
  • 3072x3072 at 20 steps also runs, but the tiled VAE decode fails there: it wants about 23 GB even when tiled.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions