Add DuCa dual feature caching as a CacheMixin.enable_cache option - #34
Open
remyx-ai[bot] wants to merge 2 commits into
Open
remyx-ai[bot] wants to merge 2 commits into
remyx-ai[bot] wants to merge 2 commits into
Conversation
Convention-shape patches extracted from huggingface/diffusers's recent merged PRs. Algorithm logic is left untouched. Ruff auto-fixed lint-trivial issues on patched files.
remyx-ai
Bot
force-pushed
the
accelerating-diffusion-transformers-with-dual-feature-cachin-v2
branch
from
September 20, 2026 14:27
ff2634f to
e7dbfd4
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Adds DuCa's training-free dual feature-caching schedule to diffusers as a new
DualCacheConfig+apply_dual_cachehook, wired into the productionCacheMixin.enable_cache/disable_cachedispatch (src/diffusers/models/cache_utils.py) so any DiT with aCacheMixincan skip block recomputation across denoising steps.enable_cachedispatches toapply_dual_cacheon aDualCacheConfig;disable_cacheremoves the hooks.What changed:
DualCachePolicyclassifies each step into compute / aggressive (reuse cached residual verbatim) / conservative (reuse with damping), with a leading full-compute step per cycle refreshing the cached full-stack residual.retention_ratio(caching disabled for initial steps) and per-cycle cache invalidation, plus dummy-object stubs and__init__.pyexports for the no-torch path.DualCacheConfigandapply_dual_cacheare exported from the package__init__.pyand reachable viapipe.transformer.enable_cache(config).Intentionally out of scope (not needed for this contribution):
conservative_scaledamping proxy because the block-level hook architecture operates on whole-block residuals and cannot selectively recompute individual tokens.Validation
Tests could not run in CI — the runner lacks this repo's dependencies (a collection/import error, not a code failure). Run the suite locally to validate.
Before submitting
Who can review?
@yiyixuxu @sayakpaul
Drafted by Outrider — paper: arXiv:2412.18911.
Discovery context
Implements Accelerating Diffusion Transformers with Dual Feature Caching.
Reference: https://github.com/Shenyi-Z/DuCa
License:
GPL-3.0(class:copyleft, compat: 0.50, source:github) — 🟡 review compatibility against this repo's license before merging.Drafted by an autonomous discovery loop — Remyx ranks recent arXiv papers against this team's research interest and shipping history; Claude Code selects the candidate most directly implementable against this repo from the lookback window and drafts it.
Research interest: [crossrepo-eval] huggingface/diffusers
Why this paper for this team: Surfaced by Outrider deep-search refine query
quantization diffusion transformer int8 fp8 inference accelerationagainst /search/assets. The engine's normal ranking did not place this paper in the interest's broad pool — it's here because the audit pass identified an under-represented theme this paper covers.Why this candidate (selected from the lookback pool): DuCa is a training-free DiT feature-caching method (cache block features at previous timesteps, reuse at next) whose I/O contract is identical to the repo's existing MagCache/TaylorSeer/FirstBlockCache hooks; it drops in as a new DualCacheConfig + apply_dual_cache branch in the already-in-production CacheMixin.enable_cache dispatch, with no new data shape or trainer required. It carries a reference implementation (GPL-3.0), lowering porting risk versus the otherwise-equivalent no-code SpeCa [12] at the same call site.
Suggested experiment: (none)
Co-Authored-By: remyx-ai[bot] <289541483+remyx-ai[bot]@users.noreply.github.com>