Skip to content

Five Z-Image pipelines miss the precomputed-t_norms DtoH-sync fix from ZImagePipeline #14879

Description

@adrianrfreedman

Describe the bug

Five Z-Image pipelines read the timestep off a device tensor inside the denoising loop, so every step blocks the CPU on a device-to-host copy:

pipeline read
z_image/pipeline_z_image_omni.py timestep[0].item()
z_image/pipeline_z_image_controlnet.py timestep[0].item()
z_image/pipeline_z_image_controlnet_inpaint.py timestep[0].item()
z_image/pipeline_z_image_img2img.py timestep[0].item()
z_image/pipeline_z_image_inpaint.py timestep[0].item()

ZImagePipeline already solves this. It hoists _precomputed_t_norms = ((1000 - timesteps.float()) / 1000).tolist() out of the loop, guarded on cfg truncation. The other five were left behind.

Same class of problem as #11696, #13404, #13406, #13461, and #13564. Found by following the examples/profiling guide, as #13401 asks.

On ZImageImg2ImgPipeline (Tongyi-MAI/Z-Image-Turbo, 1024x1024, 8 steps, one L40S) this is 1581 ms of blocked CPU per run. No wall-clock cost at this size, because the model is GPU-bound, but the loop cannot enter a CUDA graph while it syncs and the blocked CPU matters on a busier host.

Reproduction

Turn on torch.cuda.set_sync_debug_mode("error") from callback_on_step_end, after the one-off setup copies, and run any of the five. It raises RuntimeError: called a synchronizing CUDA operation at the timestep[0].item() line.

System Info

diffusers main, torch 2.9.1+cu128, one L40S.

Who can help?

@sayakpaul

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions