Skip to content

Fix missing cache contexts in Wan video-to-video - #14948

Draft
fusheng-ji wants to merge 1 commit into
huggingface:mainfrom
fusheng-ji:fix-12760-wan-v2v
Draft

fusheng-ji wants to merge 1 commit into
huggingface:mainfrom
fusheng-ji:fix-12760-wan-v2v

Conversation

@fusheng-ji

@fusheng-ji fusheng-ji commented Oct 5, 2026 •

Copy link
Copy Markdown

What does this PR do?

Enabling TaylorSeer or FirstBlockCache on WanVideoToVideoPipeline raises ValueError: No cache context is set. This wraps the conditional and unconditional transformer calls in cond and uncond cache contexts.

A shared test mixin checks CFG on/off, strength truncation, separate state per context, cleanup between calls, and unchanged output when a recording hook is attached. The change covers four files and preserves the existing transformer arguments and scheduler behavior.

Part of #12760. The maintainer acknowledged the scope and requested smaller PRs. This is the Wan video-to-video part.

Relationship to #14900

#14900 adds cache-context metadata to Wan T2V, I2V, Animate, and VACE. This PR fixes the Wan V2V path: at #14900's reviewed head (ad8168c), its conditional and unconditional transformer calls still have no named cache context.

With only the regression tests added to that head, all four V2V cases fail with No cache context is set. Applying this PR's V2V fix makes all four pass. The caching infrastructure changes remain in #14900.

Validation

Commands run from the checkout with PYTHONPATH=src:

DIFFUSERS_TEST_DEVICE=cpu OMP_NUM_THREADS=2 \
  python -m pytest tests/pipelines/wan/test_wan_video_to_video.py -k CacheContext -q
4 passed, 34 deselected

The same four cases fail with the missing-context error when run against the unchanged upstream pipeline at cff9dafbbf493e2a08c6f1951fb2d8845da71841.

DIFFUSERS_TEST_DEVICE=cpu OMP_NUM_THREADS=2 \
  python -m pytest tests/pipelines/wan/test_wan_video_to_video.py -q
DIFFUSERS_TEST_DEVICE=cuda OMP_NUM_THREADS=2 \
  python -m pytest tests/pipelines/wan/test_wan_video_to_video.py -q
CPU:  21 passed, 17 skipped
CUDA: 1 failed, 33 passed, 4 skipped

The CUDA failure is test_inference_batch_single_identical. Running that test alone on upstream reproduces the same error: 14/13056 elements outside tolerance, maximum difference 1.571774e-4, atol=1e-4, rtol=1e-5. Its tolerance and skip behavior are unchanged.

make style
make fix-copies
make quality

All three passed using the repository's quality dependencies, including ruff==0.9.10.

Real checkpoint verification

Model: Wan-AI/Wan2.1-T2V-1.3B-Diffusers, revision 0fad780a534b6463e45facd96134c9f345acfa5b. B200, torch 2.7.1+cu128, PyTorch SDPA; BF16 transformer/encoder and FP32 VAE. Fixed seed 0, CFG 5, 480×832, 33 frames; 16 requested steps with strength 0.5 give 8 denoising steps.

Check Result
Cache disabled: upstream vs fix Bitwise identical
TaylorSeer, FirstBlockCache, calibrated MagCache Complete inference with separate CFG contexts
Repeated cached calls Bitwise identical to each backend's first call
Disable cache after inference Bitwise identical to the uncached reference
FirstBlockCache reuse, threshold 0.2 Second-block computations 10/16; repeat identical
MP4 exports reread 33 frames, 8 FPS; finite output tensors

TaylorSeer uses explicit Wan attention patterns. FirstBlockCache threshold 0.05 performs full computation in this short run; threshold 0.2 exercises reuse separately. MagCache ratios are calibrated on this checkpoint with CFG disabled and used for a CFG smoke run. This verifies functionality, not cache quality or speedups. Other Wan variants and compiled execution are outside this PR.

Data and visual results

Wan V2V comparison: input, cache disabled, TaylorSeer, FirstBlockCache, and MagCache

View or download the comparison video (33 frames, 8 FPS).

Per-case results and numerical comparisons and model revision.

Images, video, and result data are stored on the fork's pr_asset branch.

Final self-review

Reviewed the complete diff against .ai/references/review-rules.md and the applicable pipeline, testing, coding-style, and numerical discrepancy guides.

  • Blocking findings: none in this patch.
  • Non-blocking findings: the existing CUDA batch/single failure above is left unchanged because it reproduces before the fix. Correcting that numerical issue would broaden this PR. Real-checkpoint coverage and cache limitations are stated above.
  • Dead code: none found. The shared mixin registers the recorder; transformer forwards enter the contexts; StateManager creates separate states; existing pipeline cleanup resets them. Tests exercise these paths and assert fresh state on the next call.
  • Documentation impact: no public API changes; existing cache usage documentation remains applicable.
  • Verdict: READY. No code fixes remain before submission. The upstream numerical failure is left for the maintainer to assess with the baseline evidence.

Before submitting

  • Used an AI agent: Codex.
  • Author has read the AI-assisted contribution guidelines.
  • Ran the self-review skill; final notes are included above.
  • Author has read the contributor guide.
  • Author has read the design philosophy.
  • Scope was discussed and acknowledged by a maintainer; coordination link is above.
  • Added regression tests and included commands, results, and baseline failures.
  • Checked documentation impact; no public API changes.
  • Images, videos, and validation artifacts are kept out of the source patch.

Model authorship is not applicable to this existing-pipeline fix.

Who can review?

@sayakpaul acknowledged the scope in the issue; review is welcome from the pipeline maintainers.

@github-actions github-actions Bot added size/M PR with diff < 200 LOC tests pipelines and removed size/M PR with diff < 200 LOC labels Oct 5, 2026
@sayakpaul

Copy link
Copy Markdown
Member

#14900 already covers Wan though.

@fusheng-ji

Copy link
Copy Markdown
Author

@sayakpaul Thanks! #14900 covers the other Wan pipelines, but this PR specifically fixes WanVideoToVideoPipeline. At its latest head (ad8168c), the V2V transformer calls still have no cache context. All four regression cases fail there and pass with this fix. Would you prefer folding the V2V fix into #14900 instead?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants