Repository navigation
Fix missing cache contexts in Wan video-to-video - #14948
Draft
fusheng-ji wants to merge 1 commit into
Draft
fusheng-ji wants to merge 1 commit into
fusheng-ji wants to merge 1 commit into
Conversation
Member
|
#14900 already covers Wan though. |
Author
|
@sayakpaul Thanks! #14900 covers the other Wan pipelines, but this PR specifically fixes WanVideoToVideoPipeline. At its latest head (ad8168c), the V2V transformer calls still have no cache context. All four regression cases fail there and pass with this fix. Would you prefer folding the V2V fix into #14900 instead? |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Enabling TaylorSeer or FirstBlockCache on
WanVideoToVideoPipelineraisesValueError: No cache context is set. This wraps the conditional and unconditional transformer calls incondanduncondcache contexts.A shared test mixin checks CFG on/off, strength truncation, separate state per context, cleanup between calls, and unchanged output when a recording hook is attached. The change covers four files and preserves the existing transformer arguments and scheduler behavior.
Part of #12760. The maintainer acknowledged the scope and requested smaller PRs. This is the Wan video-to-video part.
Relationship to #14900
#14900 adds cache-context metadata to Wan T2V, I2V, Animate, and VACE. This PR fixes the Wan V2V path: at #14900's reviewed head (
ad8168c), its conditional and unconditional transformer calls still have no named cache context.With only the regression tests added to that head, all four V2V cases fail with
No cache context is set. Applying this PR's V2V fix makes all four pass. The caching infrastructure changes remain in #14900.Validation
Commands run from the checkout with
PYTHONPATH=src:The same four cases fail with the missing-context error when run against the unchanged upstream pipeline at
cff9dafbbf493e2a08c6f1951fb2d8845da71841.The CUDA failure is
test_inference_batch_single_identical. Running that test alone on upstream reproduces the same error: 14/13056 elements outside tolerance, maximum difference1.571774e-4,atol=1e-4,rtol=1e-5. Its tolerance and skip behavior are unchanged.All three passed using the repository's quality dependencies, including
ruff==0.9.10.Real checkpoint verification
Model:
Wan-AI/Wan2.1-T2V-1.3B-Diffusers, revision0fad780a534b6463e45facd96134c9f345acfa5b. B200, torch 2.7.1+cu128, PyTorch SDPA; BF16 transformer/encoder and FP32 VAE. Fixed seed 0, CFG 5, 480×832, 33 frames; 16 requested steps with strength 0.5 give 8 denoising steps.TaylorSeer uses explicit Wan attention patterns. FirstBlockCache threshold 0.05 performs full computation in this short run; threshold 0.2 exercises reuse separately. MagCache ratios are calibrated on this checkpoint with CFG disabled and used for a CFG smoke run. This verifies functionality, not cache quality or speedups. Other Wan variants and compiled execution are outside this PR.
Data and visual results
View or download the comparison video (33 frames, 8 FPS).
Per-case results and numerical comparisons and model revision.
Images, video, and result data are stored on the fork's
pr_assetbranch.Final self-review
Reviewed the complete diff against
.ai/references/review-rules.mdand the applicable pipeline, testing, coding-style, and numerical discrepancy guides.StateManagercreates separate states; existing pipeline cleanup resets them. Tests exercise these paths and assert fresh state on the next call.Before submitting
Model authorship is not applicable to this existing-pipeline fix.
Who can review?
@sayakpaul acknowledged the scope in the issue; review is welcome from the pipeline maintainers.