Build: master-908 88411ef, Windows 11, CUDA 13.4, sm_89, RTX 4080 SUPER 16 GB, --fa --sage-attn.
Command (sd-server):
sd-server --diffusion-model qwen_image_2.1_int8_convrot.safetensors --vae qwen_image_2.1_vae_bf16.safetensors \
--llm Qwen3-VL-8B-Instruct-UD-Q4_K_XL.gguf --llm_vision mmproj-BF16.gguf --fa --sage-attn \
--model-args qwen_image_2_1_prefix_cache=false --backend te=cpu,vae=cpu
Request: edit mode, one 1792x2368 reference, output 1792x2368, euler, cfg 3, 2 steps, VAE tiling on.
Log:
[INFO ] backend_fit.cpp:348 - DiT params 6920 MiB, compute reserve 2048 MiB -> compute CUDA0, params CUDA0
[INFO ] diffusion_engine.cpp:1319 - total params memory size = 13087.89MB (VRAM 6920.55MB, RAM 6167.34MB)
[WARN ] model_manager.cpp:1665 - model manager memory on CUDA0: reported free 8133.00 MB / total 16375.50 MB, tracked weights 6920.55 MB / other runtime 0.00 MB / current runtime 0.00 MB
[WARN ] model_manager.cpp:1831 - model manager cannot make enough memory available on CUDA0: need 8133.64 MB device / 7621.64 MB budget, available 8133.00 MB device / 7628.45 MB budget
[ERROR] ggml_runner.cpp:884 - qwen_image_2_1 segment 1/1 (graph) failed during workspace capacity check
[ERROR] diffusion_engine.cpp:2551 - diffusion model compute failed
Observed: the budget check is satisfied (7621.64 needed, 7628.45 available) and the device check is missed by 0.64 MB (8133.64 needed including the 512 MiB safety margin, 8133.00 free). 6920 MB of resident, evictable weights are present. No eviction is attempted and no segmented fallback is attempted. The run is aborted.
Expected: eviction of enough resident weights to satisfy the device check, or a fallback to segmented execution when the monolithic check is missed. The same graph is executed fine in 34 segments when the weights are not resident.
Workaround (verified): --offload-to-cpu --max-vram -1. The same request completes at about 19.5 s/step.
Related, default placement: with te and vae compute left on CUDA0 (auto-fit default), about 4.5 GB is left in the CUDA VMM pool after the conditioner phase and is invisible to the model manager. The DiT run then fails at segment 10 to 12 of 34 with need 1238 MB, available 955 MB after weights have been evicted down to 2.3 GB. A negative --max-vram reserve is the only flag that accounts for it. A sentence in backend.md or performance.md saying that pool usage from earlier phases must be reserved with --max-vram -N would help.
Build: master-908
88411ef, Windows 11, CUDA 13.4, sm_89, RTX 4080 SUPER 16 GB,--fa --sage-attn.Command (sd-server):
Request: edit mode, one 1792x2368 reference, output 1792x2368, euler, cfg 3, 2 steps, VAE tiling on.
Log:
Observed: the budget check is satisfied (7621.64 needed, 7628.45 available) and the device check is missed by 0.64 MB (8133.64 needed including the 512 MiB safety margin, 8133.00 free). 6920 MB of resident, evictable weights are present. No eviction is attempted and no segmented fallback is attempted. The run is aborted.
Expected: eviction of enough resident weights to satisfy the device check, or a fallback to segmented execution when the monolithic check is missed. The same graph is executed fine in 34 segments when the weights are not resident.
Workaround (verified):
--offload-to-cpu --max-vram -1. The same request completes at about 19.5 s/step.Related, default placement: with te and vae compute left on CUDA0 (auto-fit default), about 4.5 GB is left in the CUDA VMM pool after the conditioner phase and is invisible to the model manager. The DiT run then fails at segment 10 to 12 of 34 with
need 1238 MB, available 955 MBafter weights have been evicted down to 2.3 GB. A negative--max-vramreserve is the only flag that accounts for it. A sentence in backend.md or performance.md saying that pool usage from earlier phases must be reserved with--max-vram -Nwould help.