Describe the bug
Five Z-Image pipelines read the timestep off a device tensor inside the denoising loop, so every step blocks the CPU on a device-to-host copy:
| pipeline |
read |
z_image/pipeline_z_image_omni.py |
timestep[0].item() |
z_image/pipeline_z_image_controlnet.py |
timestep[0].item() |
z_image/pipeline_z_image_controlnet_inpaint.py |
timestep[0].item() |
z_image/pipeline_z_image_img2img.py |
timestep[0].item() |
z_image/pipeline_z_image_inpaint.py |
timestep[0].item() |
ZImagePipeline already solves this. It hoists _precomputed_t_norms = ((1000 - timesteps.float()) / 1000).tolist() out of the loop, guarded on cfg truncation. The other five were left behind.
Same class of problem as #11696, #13404, #13406, #13461, and #13564. Found by following the examples/profiling guide, as #13401 asks.
On ZImageImg2ImgPipeline (Tongyi-MAI/Z-Image-Turbo, 1024x1024, 8 steps, one L40S) this is 1581 ms of blocked CPU per run. No wall-clock cost at this size, because the model is GPU-bound, but the loop cannot enter a CUDA graph while it syncs and the blocked CPU matters on a busier host.
Reproduction
Turn on torch.cuda.set_sync_debug_mode("error") from callback_on_step_end, after the one-off setup copies, and run any of the five. It raises RuntimeError: called a synchronizing CUDA operation at the timestep[0].item() line.
System Info
diffusers main, torch 2.9.1+cu128, one L40S.
Who can help?
@sayakpaul
Describe the bug
Five Z-Image pipelines read the timestep off a device tensor inside the denoising loop, so every step blocks the CPU on a device-to-host copy:
z_image/pipeline_z_image_omni.pytimestep[0].item()z_image/pipeline_z_image_controlnet.pytimestep[0].item()z_image/pipeline_z_image_controlnet_inpaint.pytimestep[0].item()z_image/pipeline_z_image_img2img.pytimestep[0].item()z_image/pipeline_z_image_inpaint.pytimestep[0].item()ZImagePipelinealready solves this. It hoists_precomputed_t_norms = ((1000 - timesteps.float()) / 1000).tolist()out of the loop, guarded on cfg truncation. The other five were left behind.Same class of problem as #11696, #13404, #13406, #13461, and #13564. Found by following the
examples/profilingguide, as #13401 asks.On
ZImageImg2ImgPipeline(Tongyi-MAI/Z-Image-Turbo, 1024x1024, 8 steps, one L40S) this is 1581 ms of blocked CPU per run. No wall-clock cost at this size, because the model is GPU-bound, but the loop cannot enter a CUDA graph while it syncs and the blocked CPU matters on a busier host.Reproduction
Turn on
torch.cuda.set_sync_debug_mode("error")fromcallback_on_step_end, after the one-off setup copies, and run any of the five. It raisesRuntimeError: called a synchronizing CUDA operationat thetimestep[0].item()line.System Info
diffusers main, torch 2.9.1+cu128, one L40S.
Who can help?
@sayakpaul