-
Notifications
You must be signed in to change notification settings - Fork 84
Pull requests: RL-Align/RL-Kernel
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
perf(rocm): MFMA batch-invariant GEMM and chunked Triton attention for strict R/R
platform: rocm
Specific tasks specific to AMD graphics cards (such as CK, bpreshuffle/FA)
fix(rocm): stabilize direct paged CK attention for bitwise R/R
platform: rocm
Specific tasks specific to AMD graphics cards (such as CK, bpreshuffle/FA)
#394
opened Sep 9, 2026 by
inaniloquentee
Collaborator
Loading…
[P5-2] Add deterministic clamp_swiglu_weighted CUDA kernel
deepseek-P5
DSv4
platform: cuda
Specific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations)
#389
opened Sep 6, 2026 by
hsiuyee
Loading…
[DSv4][P5-5] Shared Expert MLP: strict CUDA + Triton kernels
deepseek-P5
DSv4
platform: cuda
Specific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations)
#387
opened Sep 4, 2026 by
jyizheng
Loading…
[CI] Fix black formatting on main and pin line-length in pyproject.toml
#384
opened Sep 3, 2026 by
Dnoob
Contributor
Loading…
[DSv4][P1-S0] Start kit for the P1 work package (mHC + RMSNorm)
deepseek-P1
DSv4
platform: cuda
Specific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations)
#383
opened Sep 3, 2026 by
zhangj1an
Collaborator
Loading…
10 tasks
feat(ascend): add SwiGLU forward and backward kernels
Ascend
#381
opened Sep 3, 2026 by
erfgss
Contributor
Loading…
docs(readme): update feature and platform support
documentation
Improvements or additions to documentation
#374
opened Sep 1, 2026 by
Flink-ddd
Collaborator
Loading…
[WS1][Ascend] [Qwen3-8b] Fused linear logp ops
Ascend
#372
opened Sep 1, 2026 by
zhangj1an
Collaborator
Loading…
[WS1][Ascend] [Qwen3-8b] LM head ops
Ascend
#371
opened Sep 1, 2026 by
zhangj1an
Collaborator
Loading…
[WS1][Ascend] [Qwen3-8b] Fused logp ops
Ascend
#370
opened Sep 1, 2026 by
zhangj1an
Collaborator
Loading…
[WS1][Ascend] [Qwen3-8b] Embedding ops
Ascend
#369
opened Sep 1, 2026 by
zhangj1an
Collaborator
Loading…
[DSv4][P5-0] Start kit for the P5 work package (MXFP4 Routed Expert + LoRA + Shared Expert)
deepseek-P5
DSv4
platform: cuda
Specific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations)
#368
opened Sep 1, 2026 by
KJLdefeated
Collaborator
Loading…
[WS1] Add batch-invariant h_aggregate kernel
deepseek-P1
DSv4
platform: cuda
Specific optimizations or bugs in NVIDIA graphics cards (such as FlashInfer, TMA optimizations)
#366
opened Aug 30, 2026 by
nodeeeeee
Loading…
feat(ascend): add batch-invariant RMSNorm Ascend C operator
Ascend
#364
opened Aug 30, 2026 by
erfgss
Contributor
Loading…
docs: add DCO 1.1 text and contributor sign-off guide
type: ci-cd
Modify GitHub Actions, automated tests, and packaging/deployment tasks.
#359
opened Aug 29, 2026 by
Zhifu-Liu
Contributor
Loading…
feat(ascend): add deterministic collective Ascend C kernel
Ascend
#355
opened Aug 28, 2026 by
zhangj1an
Collaborator
Loading…
feat(ascend): add prefix-shared attention Ascend C kernel
Ascend
#340
opened Aug 25, 2026 by
zhangj1an
Collaborator
Loading…
[WS1][kernels] Deterministic attention Ascend C kernel
Ascend
#320
opened Aug 19, 2026 by
zhangj1an
Collaborator
Loading…
Previous Next
ProTip!
Add no:assignee to see everything that’s not assigned.