Repository navigation
perf(moe): GeGLU fused decode-MoE; wire gemma4 (+13%) (#268 step 3b) - #281
Merged
Merged
Conversation
Extend the two-kernel fused decode-MoE to GeGLU experts and route gemma4's MoE decode through it. gemma4's experts use gelu-tanh-approx(gate) * up, not SwiGLU. The gate/up kernel gains an `act` template arg (0 = SwiGLU/silu, 1 = GeGLU/gelu tanh approx, matching compiled_geglu_approx_activation); the shared two-kernel body is factored into run_fused_moe_two_kernel and exposed as two FFIs, fused_moe_expert_kernel (silu, unchanged for qwen3_moe/dots/qwen3_next) and the new fused_moe_geglu_kernel. SwitchGeGLU gains a forward_fused_kernel and Experts::forward dispatches to it on single-token decode behind MLXCEL_FUSED_MOE. gemma4's decode previously used the compiled SwitchGeGLU path (3 gather_qmm + geglu); the fused kernel reads each weight once across all cores and folds the activation + score weighting into the GEMV epilogues. Decode 73.8 -> 83.2 tok/s (+13%) on gemma-4-26b-a4b-it (M1 Ultra), the largest fused-MoE win so far (small experts, Dff=704, make the MoE a meaningful decode fraction). Validated greedy temp-0 byte-identical to the compiled-switch path with the chat template on gemma-4-26b-a4b-it; 4-bit affine, off by default. fmt + clippy clean.
This was referenced Jun 14, 2026
inureyes
added a commit
that referenced
this pull request
Jun 14, 2026
…#283) * docs(moe): document MLXCEL_FUSED_MOE flags, per-model gains, and M5 follow-up (#268) Add a usage/flags section to the fused decode-MoE design doc: the MLXCEL_FUSED_MOE / MLXCEL_FUSED_MOE_SGY / MLXCEL_FUSED_MOE_RELU2 env vars, the measured per-model decode gains on M1 Ultra (gemma4 +13%, qwen3.5 +8.7%, dots +4.7%, qwen3-30b +3.5%, nemotron ~0%), the f16-jitter-class parity caveat, and the list of covered models. Mark roadmap steps 3-4 done (#278/#279/#280/#281) and add step 5: validate on M5, then decide on flipping the flag default-on. * docs(moe): reference the M5 validation issue (#282)
This was referenced Jun 14, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Step 3b of #268: extend the two-kernel fused decode-MoE to GeGLU experts and route gemma4's MoE decode through it. gemma4's experts use
gelu-tanh-approx(gate) * up, not SwiGLU.What changed
acttemplate arg: 0 = SwiGLU (silu), 1 = GeGLU (gelu tanh approx, matchingcompiled_geglu_approx_activation).run_fused_moe_two_kernel, exposed as two FFIs:fused_moe_expert_kernel(silu — unchanged for qwen3_moe/dots/qwen3_next) and the newfused_moe_geglu_kernel.SwitchGeGLUgainsforward_fused_kernel;Experts::forwarddispatches to it on single-token decode behindMLXCEL_FUSED_MOE.Result (gemma-4-26b-a4b-it, M1 Ultra)
Decode 73.8 → 83.2 tok/s (+13%) — the largest fused-MoE win so far. gemma4 previously used the compiled SwitchGeGLU path (3
gather_qmm+ geglu); the fused kernel reads each weight once across all cores and folds the activation + score weighting into the GEMV epilogues. Its small experts (Dff=704) make the MoE a meaningful decode fraction.Validation
MLXCEL_FUSED_MOE).cargo fmt+cargo clippyclean.Refs #268.