Skip to content

hires with a model-backed upscaler aborts on Qwen-Image-2.1 (GGML_ASSERT ne[2] in ggml_im2col) #2027

Description

@nbeerbower

What happens

With Qwen-Image-2.1 loaded, an img_gen request whose hires.upscaler is a
model-backed upscaler (ESRGAN) aborts the process:

[INFO ] image.cpp:913  - hires fix: upscaling to 1024x1024
[INFO ] image.cpp:921  - hires fix: loading model upscaler from '.../RealESRGAN_x4plus_anime_6B.pth'
[INFO ] image.cpp:945  - hires fix: scheduler_steps=4, denoising_strength=0.50, sigma_sched_size=4
ggml/src/ggml.c:4557: GGML_ASSERT(a->ne[2] == b->ne[2]) failed

That assertion is the is_2D branch of ggml_im2col. The server does not
return an error, it dies, so every other queued job goes with it.

A latent upscaler on the same model is fine, which is what makes this look
like a shape mismatch rather than a bad model file.

Reproduction

Same server, same request, only hires.upscaler differs.

Server:

sd-server --diffusion-model qwen_image_2.1-Q8_0.gguf \
          --vae qwen_image_2.1_vae_bf16.safetensors \
          --llm Qwen3VL-8B-Instruct-Q8_0.gguf \
          --llm_vision mmproj-Qwen3VL-8B-Instruct-F16.gguf \
          --hires-upscalers-dir <dir with RealESRGAN_x4plus_anime_6B.pth> \
          --diffusion-fa

Works — completes in about 7 s:

{"prompt":"a cat","width":512,"height":512,"seed":1,
 "sample_params":{"sample_method":"euler","sample_steps":4},
 "hires":{"enabled":true,"upscaler":"Latent","scale":2.0,
          "steps":2,"denoising_strength":0.5}}

Aborts the server:

{"prompt":"a cat","width":512,"height":512,"seed":1,
 "sample_params":{"sample_method":"euler","sample_steps":4},
 "hires":{"enabled":true,"upscaler":"RealESRGAN_x4plus_anime_6B","scale":2.0,
          "steps":2,"denoising_strength":0.5}}

Guess at the cause

I have not fixed this, so this is only where I would start looking. The two
paths diverge at image.cpp:708: a model-backed upscaler decodes the latent
to an image, runs the upscaler, and re-encodes, where a latent upscaler stays
in latent space. Qwen Image 2.1 carries qwen_image_layers + 1 in the layer
dimension (image.cpp:221, generate_init_latent(..., request->qwen_image_layers + 1, true),
default 3, so 4), and a latent re-encoded from a plain RGB image would have 1
there. That would fit an ne[2] mismatch reaching ggml_im2col.

Environment

  • master at 6dcb5bb, CUDA build, RTX A6000, Linux
  • Qwen-Image-2.1 Q8_0 (leejet/Qwen-Image-2.1-GGUF) with
    qwen_image_2.1_vae_bf16.safetensors and Qwen3-VL-8B-Instruct Q8_0
  • RealESRGAN_x4plus_anime_6B.pth

Standalone upscaling with the same model file is fine (sd-cli -M upscale,
512x512 to 2048x2048 in ~2 s), so the upscaler itself loads and runs.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions