Skip to content

[Bug] Speed is halved during qwen image 2.1 reference generation #2031

Description

@brt16

Git commit

c92d73c

Operating System & Version

windows 11

GGML backends

Vulkan

Command-line arguments used

I used the UI but it should be something like qwen_image_2.1-Q8_0.gguf, --diffusion-fa --offload-to-cpu, euler, seed 42

Steps to reproduce

generating an image with references

What you expected to happen

There should be an option to not process references at every step

What actually happened

It processes the references at every step, doubling generation time while there are reference images

Logs / error messages / stack trace

No response

Additional context / environment details

AI generated fix is tested, outputs are byte-identical and should work perfectly. It should ideally be made an option, as it expends more VRAM in return, but it works faster even with ram spill, so it's a no brainer.

qwen-image-2.1-prefix-cache.patch

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions