Git commit
c92d73c
Operating System & Version
windows 11
GGML backends
Vulkan
Command-line arguments used
I used the UI but it should be something like qwen_image_2.1-Q8_0.gguf, --diffusion-fa --offload-to-cpu, euler, seed 42
Steps to reproduce
generating an image with references
What you expected to happen
There should be an option to not process references at every step
What actually happened
It processes the references at every step, doubling generation time while there are reference images
Logs / error messages / stack trace
No response
Additional context / environment details
AI generated fix is tested, outputs are byte-identical and should work perfectly. It should ideally be made an option, as it expends more VRAM in return, but it works faster even with ram spill, so it's a no brainer.
qwen-image-2.1-prefix-cache.patch
Git commit
c92d73c
Operating System & Version
windows 11
GGML backends
Vulkan
Command-line arguments used
I used the UI but it should be something like
qwen_image_2.1-Q8_0.gguf,--diffusion-fa --offload-to-cpu, euler, seed 42Steps to reproduce
generating an image with references
What you expected to happen
There should be an option to not process references at every step
What actually happened
It processes the references at every step, doubling generation time while there are reference images
Logs / error messages / stack trace
No response
Additional context / environment details
AI generated fix is tested, outputs are byte-identical and should work perfectly. It should ideally be made an option, as it expends more VRAM in return, but it works faster even with ram spill, so it's a no brainer.
qwen-image-2.1-prefix-cache.patch