Repository navigation
Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
Contributor
This was referenced Sep 23, 2026
This was referenced Sep 23, 2026
kaix-nv
added this pull request to stack #2521
September 23, 2026 01:22
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #2519 +/- ##
==========================================
+ Coverage 71.76% 78.11% +6.34%
==========================================
Files 644 659 +15
Lines 71477 72634 +1157
==========================================
+ Hits 51298 56740 +5442
+ Misses 20179 15894 -4285
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
kaix-nv
force-pushed
the
kaix/linear-attention-decode-first
branch
from
September 23, 2026 19:34
94080d3 to
757f337
Compare
kaix-nv
removed this pull request from stack #2521
September 24, 2026 06:17
kaix-nv
added this pull request to stack #2542
September 24, 2026 06:18
kaix-nv
removed this pull request from stack #2542
September 24, 2026 06:31
kaix-nv
added this pull request to stack #2543
September 24, 2026 06:31
kaix-nv
force-pushed
the
kaix/linear-attention-decode-first
branch
from
September 24, 2026 18:17
757f337 to
492db57
Compare
kaix-nv
force-pushed
the
kaix/linear-attention-decode-first
branch
from
September 25, 2026 01:54
492db57 to
5fbb898
Compare
kaix-nv
force-pushed
the
kaix/linear-attention-decode-first
branch
3 times, most recently
from
September 25, 2026 20:57
5e548c1 to
ab35f1e
Compare
kaix-nv
force-pushed
the
kaix/linear-attention-decode-first
branch
4 times, most recently
from
September 28, 2026 06:05
30e0659 to
7b5caf1
Compare
kaix-nv
removed this pull request from stack #2658
October 8, 2026 18:41
kaix-nv
added this pull request to stack #2714
October 8, 2026 18:41
kaix-nv
force-pushed
the
kaix/linear-attention-decode-first
branch
from
October 8, 2026 22:40
94c5ead to
2c90acd
Compare
kaix-nv
added a commit
that referenced
this pull request
Oct 10, 2026
Consolidate the unpublished PR #2519 review fixes. Preserve valid packed lengths through selective recompute, expose the public phase API, and keep the Bridge workflow integration separate. Retain the existing FLA kernels for their follow-up PR. Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Use registered state quantizers for format and last-axis grouping, with INT8 groups of 16, 32, or 64 independent of execution tile width. Preserve existing tile and Hadamard codecs and document the blockwise configuration. Keep minimal output/gradient, configuration, and checkpoint coverage; omit redundant mocked block routing and quantizer call-count checks. Signed-off-by: Kai Xu <kaix@nvidia.com>
Make recurrent_decode the single Torch implementation. Process each prepared prefix directly and remove duplicate packing, shape preparation, and the obsolete helper without changing supported numerical policies. Validated 40 focused CPU tests, the existing Bridge QAT/QAD GPU smoke tests, and 12 before/after output, state, and gradient comparisons. Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Unify prefix, decode, and replay execution settings in LinearAttentionConfig with migration for supported nested configs and saved policy objects. Keep formats and grouping in TensorQuantizer. Retire W-only and reference training paths, move mathematical references under tests, and remove the copied FLA W-QAT implementation and unused recipes. Validation: 44 focused CPU tests, 6 native and Megatron GPU tests, legacy checkpoint loading, recipe validation, and pre-commit hooks passed. Signed-off-by: Kai Xu <kaix@nvidia.com>
Normalize recurrent working values to FP32 and require an explicit serving policy for both GDN and KDA state quantization. Remove unpublished config migrations, unused replay-factor plumbing, and the unreferenced Triton INT8 helper while retaining native INT8/Hadamard replay. Align GDN/KDA API and runtime-state names, document the state handoff, and keep shared TensorQuantizer/FP8 corrections outside this PR. Update recipe guidance and retain minimal real-path tests. Validation: 28 focused CPU tests, 5 native GPU cases, 2 Megatron QAT/sharded-restore tests, scoped pre-commit hooks, and git diff --check passed. Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Adapt native imports, state layouts, recurrent indexing, and KDA gate arithmetic while keeping ModelOpt state QDQ in canonical layout. Use the vllm profile name and retain vllm_0_15 as a legacy alias. Keep lazy gate discovery outside compiled graphs. Validate policy matches across distributed stages and correct the optional dependency message. Verified 94 CPU tests, native GDN/KDA GPU checks on vLLM 0.15.1, 0.20.0, and 0.30.0, ReplaySSM cases, Megatron backward/checkpoint restore, and pre-commit hooks. Full serving-engine and model-quality qualification remain separate. Signed-off-by: Kai Xu <kaix@nvidia.com>
Cache state-layout and kernel-signature checks at first use so Sphinx can import serving adapters with mocked optional dependencies. Preserve runtime dispatch and kernel arithmetic. Validation: focused recursive Sphinx HTML build using repository configuration and warnings as errors; native metadata and CPU state-layout checks on vLLM 0.15.1, 0.20.0, and 0.30.0; scoped pre-commit and git diff --check. Signed-off-by: Kai Xu <kaix@nvidia.com>
Consolidate the unpublished PR #2519 review fixes. Preserve valid packed lengths through selective recompute, expose the public phase API, and keep the Bridge workflow integration separate. Retain the existing FLA kernels for their follow-up PR. Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
kaix-nv
force-pushed
the
kaix/linear-attention-decode-first
branch
from
October 10, 2026 04:40
fec406f to
140735b
Compare
…lation Signed-off-by: Kai Xu <kaix@nvidia.com>
Signed-off-by: Kai Xu <kaix@nvidia.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Linear-attention series — 6 PRs
mainThe five open PRs form one native GitHub stack in the order shown. #2497 has landed, so #2519 targets
main. #2541 now targets #2657. Rebase each remaining descendant after its immediate parent merges.#2541 applies TensorQuantizer before native vLLM prefill/decode calls. Serving-time prefill-GEMM quantization remains deferred until an optimized fused kernel is available.
What does this PR do?
Type of change: New feature.
Add recurrent-state QAT for Megatron-Core GDN and KDA. Each sequence uses a chunked prefill prefix followed by a recurrent suffix. TensorQuantizer applies fake QDQ at the handoff and subsequent state writes, so training includes the recurrent quantization effects encountered during decode.
An explicit linear-attention policy selects the native forward arithmetic and state/checkpoint schedule. Training uses native forward values with differentiable FLA/Torch surrogate gradients and identity STE through QDQ; this does not claim an exact backward for inference kernels.
precision="vllm"selects the supported installed Triton APIs. GDN qualification covers pure single-token decode with matching Triton prefill and packed-decode settings. Generic KDA selects standalone FLA/Triton APIs without model-level serving parity.precision="vllm_kimi_k3"selects the Kimi-K3/Kimi-Linear Triton profile. Serving must use Triton prefill/decode, Q/K normalization andgate_lower_bound=None.precision="replayssm"requires an optional private quantized-ReplaySSM fork, unavailable in public vLLM releases. Without access to it, use a block32 recipe. The fork profile uses Hadamard checkpoints withreplay_window=1by default.The public training phase supplies per-sequence prefix and valid-token lengths, including packed padding. Active state QAT requires native Megatron
_compute_gatesand_forward_computehooks for raw gates and packed sequence lengths; Core 0.19.2 lacks them and is rejected before forward execution. Selective GDN recompute captures the phase for backward on supported Core builds. Active state QAT rejects context parallelism and full-layer recompute. Disabled state quantizers outside a phase preserve the original module path.This change supports floating-storage state fake quantization. It does not add a vLLM server or native compressed cache. #2541 owns serving integration; #2657 owns the training example. Prefill GEMM quantization remains a follow-up. Descendant rebases remain deferred until their immediate parent merges.
Usage
This expects an initialized Megatron model and a loss on the decode suffix. Training needs compatible vLLM and FLA packages. Validation/calibration loops must also establish the appropriate phase. A full-prefill pass does not measure decode state-QDQ quality. Compare against a converted model with the same serving policy and disabled state quantizers.
Testing
Current head:
03f6076dd8. CI results for this head are pending. The prior GPU CI run on140735b962had these results:test_mamba_search_spacetimed out; KDA skips because the image lacks its module.Local validation of the code published in
03f6076dd8, on RTX A6000:ac100f773f9d/ Torch 2.9.1 / Transformer Engine 2.16 / vLLM 0.15.2.dev: 2 GDN tests passed, including QAT/sharded restore, packed recompute and the raw-gate/alternate-entry guards.git diff --check: passed.The runtime rejects sm_120+ with Triton < 3.7 before state-QAT kernels launch. The whole training test module also skips this combination, including layout probes. The observed misaligned-address fault has not been isolated to a specific kernel; it affects real training. vLLM 0.20/sm_120 therefore provides no state-training coverage with this mitigation; its new CI result is pending.
The suite retains native comparisons and shares K=128 between generic and Kimi KDA to reuse compilation. The workflow raises only the vLLM 0.30 timeout from 15 to 25 minutes. Setup CODEOWNER review and a green CI run are still required.
KDA bit-exact qualification covers the tested vLLM 0.30 kernel profiles/shapes only. This does not establish full-engine scheduling parity, model-quality recovery or training speed. The private ReplaySSM profile remains outside public CI coverage.
Before your PR is "Ready for review"
fla-core==0.5.1without replacing vLLM's TileLang/TVM pins.Additional Information
Step 2/6. Known limits include per-layer boundary synchronization, an additional FLA prefix forward for gradients, per-token native cache allocations, and explicit valid lengths for padded-only packed metadata. The source size is intentionally deferred for this review.
Summary by CodeRabbit