Describe the bug
Three Cosmos pipelines read the timestep off a device tensor inside the denoising loop, so every step blocks the CPU on a device-to-host copy:
| pipeline |
read |
cosmos/pipeline_cosmos2_5_predict.py |
t.cpu().item() |
cosmos/pipeline_cosmos2_5_transfer.py |
t.cpu().item() |
cosmos/pipeline_cosmos3_omni.py |
t.item() |
cosmos3_omni has a second one. _mask_velocity_predictions guards each modality with if mask.sum() > 0, a device reduction read on the host once per modality per pass per step. The guards matter and cannot simply be dropped, because 0 * inf and 0 * nan are both NaN, so an all-zero mask with a non-finite prediction gives zeros today and would start giving NaN without them. The condition masks are fixed for a whole run, so the flags only need reading once.
Same class of problem as #11696, #13404, #13406, #13461, and #13564. Found by following the examples/profiling guide, as #13401 asks.
On Cosmos3OmniPipeline (nvidia/Cosmos3-Edge, 6 steps, one L40S) the loop makes 27 device-to-host copies and blocks the CPU for 13.23 s per run.
Cosmos3 will not be sync-free even once the loop is fixed. UniPCMultistepScheduler keeps self.sigmas on the CPU and copies a scalar to the device four times in scheduling_unipc_multistep.py, once per step. UniPC is used by 20 pipelines, so that belongs in its own issue.
Reproduction
Profile Cosmos3OmniPipeline with torch.profiler and count the Memcpy DtoH events inside the denoising loop, or turn on torch.cuda.set_sync_debug_mode("error") from a step callback after the one-off setup copies.
System Info
diffusers main, torch 2.9.1+cu128, one L40S.
Who can help?
@sayakpaul
Describe the bug
Three Cosmos pipelines read the timestep off a device tensor inside the denoising loop, so every step blocks the CPU on a device-to-host copy:
cosmos/pipeline_cosmos2_5_predict.pyt.cpu().item()cosmos/pipeline_cosmos2_5_transfer.pyt.cpu().item()cosmos/pipeline_cosmos3_omni.pyt.item()cosmos3_omnihas a second one._mask_velocity_predictionsguards each modality withif mask.sum() > 0, a device reduction read on the host once per modality per pass per step. The guards matter and cannot simply be dropped, because0 * infand0 * nanare both NaN, so an all-zero mask with a non-finite prediction gives zeros today and would start giving NaN without them. The condition masks are fixed for a whole run, so the flags only need reading once.Same class of problem as #11696, #13404, #13406, #13461, and #13564. Found by following the
examples/profilingguide, as #13401 asks.On
Cosmos3OmniPipeline(nvidia/Cosmos3-Edge, 6 steps, one L40S) the loop makes 27 device-to-host copies and blocks the CPU for 13.23 s per run.Cosmos3 will not be sync-free even once the loop is fixed.
UniPCMultistepSchedulerkeepsself.sigmason the CPU and copies a scalar to the device four times inscheduling_unipc_multistep.py, once per step. UniPC is used by 20 pipelines, so that belongs in its own issue.Reproduction
Profile
Cosmos3OmniPipelinewithtorch.profilerand count theMemcpy DtoHevents inside the denoising loop, or turn ontorch.cuda.set_sync_debug_mode("error")from a step callback after the one-off setup copies.System Info
diffusers main, torch 2.9.1+cu128, one L40S.
Who can help?
@sayakpaul