Skip to content

CUDA: monolithic workspace check missed by 0.64 MB with all weights resident; no eviction or segmented fallback (Qwen-Image-2.1 edit, 1792x2368, 16 GB) #2042

Description

@SixSeven-Labs

Build: master-908 88411ef, Windows 11, CUDA 13.4, sm_89, RTX 4080 SUPER 16 GB, --fa --sage-attn.

Command (sd-server):

sd-server --diffusion-model qwen_image_2.1_int8_convrot.safetensors --vae qwen_image_2.1_vae_bf16.safetensors \
  --llm Qwen3-VL-8B-Instruct-UD-Q4_K_XL.gguf --llm_vision mmproj-BF16.gguf --fa --sage-attn \
  --model-args qwen_image_2_1_prefix_cache=false --backend te=cpu,vae=cpu

Request: edit mode, one 1792x2368 reference, output 1792x2368, euler, cfg 3, 2 steps, VAE tiling on.

Log:

[INFO ] backend_fit.cpp:348 - DiT params 6920 MiB, compute reserve 2048 MiB -> compute CUDA0, params CUDA0
[INFO ] diffusion_engine.cpp:1319 - total params memory size = 13087.89MB (VRAM 6920.55MB, RAM 6167.34MB)
[WARN ] model_manager.cpp:1665 - model manager memory on CUDA0: reported free 8133.00 MB / total 16375.50 MB, tracked weights 6920.55 MB / other runtime 0.00 MB / current runtime 0.00 MB
[WARN ] model_manager.cpp:1831 - model manager cannot make enough memory available on CUDA0: need 8133.64 MB device / 7621.64 MB budget, available 8133.00 MB device / 7628.45 MB budget
[ERROR] ggml_runner.cpp:884  - qwen_image_2_1 segment 1/1 (graph) failed during workspace capacity check
[ERROR] diffusion_engine.cpp:2551 - diffusion model compute failed

Observed: the budget check is satisfied (7621.64 needed, 7628.45 available) and the device check is missed by 0.64 MB (8133.64 needed including the 512 MiB safety margin, 8133.00 free). 6920 MB of resident, evictable weights are present. No eviction is attempted and no segmented fallback is attempted. The run is aborted.

Expected: eviction of enough resident weights to satisfy the device check, or a fallback to segmented execution when the monolithic check is missed. The same graph is executed fine in 34 segments when the weights are not resident.

Workaround (verified): --offload-to-cpu --max-vram -1. The same request completes at about 19.5 s/step.

Related, default placement: with te and vae compute left on CUDA0 (auto-fit default), about 4.5 GB is left in the CUDA VMM pool after the conditioner phase and is invisible to the model manager. The DiT run then fails at segment 10 to 12 of 34 with need 1238 MB, available 955 MB after weights have been evicted down to 2.3 GB. A negative --max-vram reserve is the only flag that accounts for it. A sentence in backend.md or performance.md saying that pool usage from earlier phases must be reserved with --max-vram -N would help.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions