Repository navigation
Track and root-cause the 23 GB10 CUDA model failures at 0.3.0 #315
Description
Activity
- addedstatus:investigationFeasibility spike / under investigationFeasibility spike / under investigationtype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionspriority:highHigh priorityHigh priorityarea:benchmarkBenchmark harness and performance measurement (bench_*.sh, /update-benchmarks)Benchmark harness and performance measurement (bench_*.sh, /update-benchmarks)area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)
on Jun 16, 2026 Update: the fused-MoE-path CUDA failures are fixed by #319
Re-benchmarking on latest
main(the #307 wiring epic in progress) showed the fused decode-MoE kernel does not just leave the Qwen MoE broken on CUDA, it newly broke two more models that previously passed on the legacy path:lfm2-8b-a1b-4bit(wired by #309) andqwen3-vl-30b-a3b-4bit(wired by #310).mixtralandphimoe(wired by #311/#312) happen to survive.#319 defaults the fused decode-MoE kernel off on CUDA. Post-fix GB10 re-bench (no env override) restores all seven fused-MoE-path models to their pre-regression decode rates:
Model after #319 qwen3-moe-4bit 58.88 qwen3-30b-a3b-4bit 57.31 qwen3.5-35b-a3b-4bit 46.77 qwen1.5-moe-a2.7b-4bit 112.29 qwen3.6-35b-a3b-4bit 45.25 lfm2-8b-a1b-4bit 131.88 qwen3-vl-30b-a3b-4bit 57.50 Remaining items in this catalog (not addressed by #319)
- Other MoE (
gemma-4-26b-a4b-it-4bit,gemma-4-26b-a4b-it-qat-4bit,deepseek-v3-4bit,dots.llm1.inst-mixed-4-6bit): these use a different MoE path (not the fused gate) and still fail; need separate investigation. - Non-MoE (
bitnetx2,glm-5/glm-5.1,paligemma2-3b-6bit0-token): unchanged, still open. - By-design (drafters, non-text-gen, VLM variants): reclassify/exclude.
- Other MoE (
Update: #319 was repointed from the CUDA fused-off workaround to a CUDA port of the fused decode-MoE kernel (greedy byte-identical, +10% to +55% on GB10). The seven fused-MoE-path models are now fixed and faster, not just falling back. The remaining non-fused failures in this catalog (gemma-4-26b-a4b, deepseek-v3, dots.llm1, glm-5, bitnet, paligemma2) are unaffected and still open.
- addedstatus:doneCompletedCompletedand removedstatus:investigationFeasibility spike / under investigationFeasibility spike / under investigation
on Sep 6, 2026
Summary
The 2026-06-16 GB10 (DGX Spark) full benchmark at mlxcel 0.3.0 (CUDA 13.0, SM 12.1) recorded 124 pass / 23 fail / 1 OOM-skip across 148 text models. This issue tracks the root cause of each of the 23 failures and classifies them as genuine bugs to fix vs by-design exclusions to reclassify. Data:
benchmarks/cuda_gb10_2026-06-16.csv; per-family breakdown indocs/benchmark_results/model_tests_gb10.md. All 23 areFAIL:bench(non-OOM warmup/bench failure) exceptpaligemma2-3b-6bit, which loads and prefills but generates 0 tokens.Reproduce any single model:
./scripts/bench_decode.sh models/<name>or./target/release/mlxcel-bench-decode -m models/<name> -p "Hello" -n 20 --warmup-tokens 5.Genuine failures to fix
Qwen fused-MoE regression — tracked in #313 (5)
qwen1.5-moe-a2.7b-4bit,qwen3-moe-4bit,qwen3-30b-a3b-4bit,qwen3.5-35b-a3b-4bit,qwen3.6-35b-a3b-4bitRegression vs 0.1.0 on the
qwen2_moe/qwen3_moefused-MoE CUDA path. Tracked in #313 (with the #307 wiring epic); not re-investigated here.Other MoE warmup failures (4) — likely the same fused-MoE path
gemma-4-26b-a4b-it-4bit,gemma-4-26b-a4b-it-qat-4bit— Gemma-4 A4B MoE. Also failed at 0.1.0 (not a 0.3.0 regression), but check whether the fused-MoE wiring / CUDA fallback (GB10 CUDA: Qwen fused-MoE models fail warmup at 0.3.0 (regression vs 0.1.0) #313, epic(moe): wire the remaining unwired MoE families to the fused decode-MoE kernel #307) covers them.dots.llm1.inst-mixed-4-6bit— dots.llm1 MoE (mixed 4/6-bit). Check loader + MoE path on CUDA.deepseek-v3-4bit— very large MoE; also failed at 0.1.0. Confirm whether this is a capacity issue (weights vs the 122 GB budget) or an architecture/kernel failure.BitNet ternary (2)
bitnet-b1.58-2b-4t,bitnet-b1.58-2b-4t-4bit— ternary BitNet. Determine whether the CUDA backend supports the BitNet quant/kernels; if unsupported, document it, otherwise close the kernel gap.GLM-5 (2)
glm-5-4bit,glm-5.1-4bit— GLM-5 family. Capture the warmup error and check the GLM-5 loader/arch on CUDA. (glm4-flash-4bitpasses at 53.94 tok/s, so this is GLM-5-specific.)Zero tokens generated (1)
paligemma2-3b-6bit— loads and prefills (163.57 tok/s) but emits 0 decode tokens. Immediate-EOS / sampling / generation-loop issue, not a load failure (same 0-token behavior at 0.1.0).Likely by-design — reclassify / exclude from the suite (not bugs)
Not standalone text-generation models (3)
docling-layout-heron-mlx-bf16(document layout),granite-speech-4.1-2b-nar-mlx(speech),diffusiongemma-26b-a4b-it-4bit(diffusion LM, non-autoregressive)These are not autoregressive text generators, so the decode bench cannot drive them. Consider excluding them from the text suite or marking them non-standalone.
MTP / DFlash drafter checkpoints (4) — need a target model
gemma-4-12b-it-assistant-4bit,gemma-4-31b-it-assistant-bf16(MTP assistant drafters),qwen3.5-27b-dflash,qwen3.5-4b-dflash(DFlash drafters)Drafters bind to a target for speculative decoding and are not standalone. Exclude from the standalone bench or document.
VLM variants under the text/image setup (2)
qwen2.5-vl-3b(bf16) — fails the text-prompt pass (also failed at 0.1.0). Noteqwen2.5-vl-3b-4bitworks via the image path (60.24 tok/s), so this is the bf16 variant specifically.minicpm-v-4.6-mxfp4— the mxfp4 variant fails;minicpm-v-4.6-bf16passes (110.67 tok/s). Likely an mxfp4 quant path issue on CUDA.Notes
qwen3-next-480b-4bit, weights exceed the 122 GB budget) is a capacity exclusion, not a failure, and is out of scope.docs/benchmark_results/model_tests_m1ultra.md/model_tests_m5max.mdto separate CUDA-specific failures from universal ones.Counts: genuine-to-fix = 14 (5 in #313 + 4 MoE + 2 BitNet + 2 GLM-5 + 1 zero-token), by-design = 9 (3 non-text-gen + 4 drafters + 2 VLM). Total 23.