Skip to content

Track and root-cause the 23 GB10 CUDA model failures at 0.3.0 #315

Description

@inureyes

Summary

The 2026-06-16 GB10 (DGX Spark) full benchmark at mlxcel 0.3.0 (CUDA 13.0, SM 12.1) recorded 124 pass / 23 fail / 1 OOM-skip across 148 text models. This issue tracks the root cause of each of the 23 failures and classifies them as genuine bugs to fix vs by-design exclusions to reclassify. Data: benchmarks/cuda_gb10_2026-06-16.csv; per-family breakdown in docs/benchmark_results/model_tests_gb10.md. All 23 are FAIL:bench (non-OOM warmup/bench failure) except paligemma2-3b-6bit, which loads and prefills but generates 0 tokens.

Reproduce any single model: ./scripts/bench_decode.sh models/<name> or ./target/release/mlxcel-bench-decode -m models/<name> -p "Hello" -n 20 --warmup-tokens 5.

Genuine failures to fix

Qwen fused-MoE regression — tracked in #313 (5)

  • qwen1.5-moe-a2.7b-4bit, qwen3-moe-4bit, qwen3-30b-a3b-4bit, qwen3.5-35b-a3b-4bit, qwen3.6-35b-a3b-4bit

Regression vs 0.1.0 on the qwen2_moe / qwen3_moe fused-MoE CUDA path. Tracked in #313 (with the #307 wiring epic); not re-investigated here.

Other MoE warmup failures (4) — likely the same fused-MoE path

BitNet ternary (2)

  • bitnet-b1.58-2b-4t, bitnet-b1.58-2b-4t-4bit — ternary BitNet. Determine whether the CUDA backend supports the BitNet quant/kernels; if unsupported, document it, otherwise close the kernel gap.

GLM-5 (2)

  • glm-5-4bit, glm-5.1-4bit — GLM-5 family. Capture the warmup error and check the GLM-5 loader/arch on CUDA. (glm4-flash-4bit passes at 53.94 tok/s, so this is GLM-5-specific.)

Zero tokens generated (1)

  • paligemma2-3b-6bit — loads and prefills (163.57 tok/s) but emits 0 decode tokens. Immediate-EOS / sampling / generation-loop issue, not a load failure (same 0-token behavior at 0.1.0).

Likely by-design — reclassify / exclude from the suite (not bugs)

Not standalone text-generation models (3)

  • docling-layout-heron-mlx-bf16 (document layout), granite-speech-4.1-2b-nar-mlx (speech), diffusiongemma-26b-a4b-it-4bit (diffusion LM, non-autoregressive)

These are not autoregressive text generators, so the decode bench cannot drive them. Consider excluding them from the text suite or marking them non-standalone.

MTP / DFlash drafter checkpoints (4) — need a target model

  • gemma-4-12b-it-assistant-4bit, gemma-4-31b-it-assistant-bf16 (MTP assistant drafters), qwen3.5-27b-dflash, qwen3.5-4b-dflash (DFlash drafters)

Drafters bind to a target for speculative decoding and are not standalone. Exclude from the standalone bench or document.

VLM variants under the text/image setup (2)

  • qwen2.5-vl-3b (bf16) — fails the text-prompt pass (also failed at 0.1.0). Note qwen2.5-vl-3b-4bit works via the image path (60.24 tok/s), so this is the bf16 variant specifically.
  • minicpm-v-4.6-mxfp4 — the mxfp4 variant fails; minicpm-v-4.6-bf16 passes (110.67 tok/s). Likely an mxfp4 quant path issue on CUDA.

Notes

  • The OOM-skip (qwen3-next-480b-4bit, weights exceed the 122 GB budget) is a capacity exclusion, not a failure, and is out of scope.
  • Cross-check Apple Silicon status in docs/benchmark_results/model_tests_m1ultra.md / model_tests_m5max.md to separate CUDA-specific failures from universal ones.

Counts: genuine-to-fix = 14 (5 in #313 + 4 MoE + 2 BitNet + 2 GLM-5 + 1 zero-token), by-design = 9 (3 non-text-gen + 4 drafters + 2 VLM). Total 23.

Activity

  1. added
    type:bugBug fixes, error corrections, or issue resolutions
    area:benchmarkBenchmark harness and performance measurement (bench_*.sh, /update-benchmarks)
    area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)
    on Jun 16, 2026
  2. inureyes commented on Jun 17, 2026

    @inureyes
    MemberAuthor

    Update: the fused-MoE-path CUDA failures are fixed by #319

    Re-benchmarking on latest main (the #307 wiring epic in progress) showed the fused decode-MoE kernel does not just leave the Qwen MoE broken on CUDA, it newly broke two more models that previously passed on the legacy path: lfm2-8b-a1b-4bit (wired by #309) and qwen3-vl-30b-a3b-4bit (wired by #310). mixtral and phimoe (wired by #311/#312) happen to survive.

    #319 defaults the fused decode-MoE kernel off on CUDA. Post-fix GB10 re-bench (no env override) restores all seven fused-MoE-path models to their pre-regression decode rates:

    Model after #319
    qwen3-moe-4bit 58.88
    qwen3-30b-a3b-4bit 57.31
    qwen3.5-35b-a3b-4bit 46.77
    qwen1.5-moe-a2.7b-4bit 112.29
    qwen3.6-35b-a3b-4bit 45.25
    lfm2-8b-a1b-4bit 131.88
    qwen3-vl-30b-a3b-4bit 57.50

    Remaining items in this catalog (not addressed by #319)

    • Other MoE (gemma-4-26b-a4b-it-4bit, gemma-4-26b-a4b-it-qat-4bit, deepseek-v3-4bit, dots.llm1.inst-mixed-4-6bit): these use a different MoE path (not the fused gate) and still fail; need separate investigation.
    • Non-MoE (bitnet x2, glm-5/glm-5.1, paligemma2-3b-6bit 0-token): unchanged, still open.
    • By-design (drafters, non-text-gen, VLM variants): reclassify/exclude.
  3. inureyes commented on Jun 17, 2026

    @inureyes
    MemberAuthor

    Update: #319 was repointed from the CUDA fused-off workaround to a CUDA port of the fused decode-MoE kernel (greedy byte-identical, +10% to +55% on GB10). The seven fused-MoE-path models are now fixed and faster, not just falling back. The remaining non-fused failures in this catalog (gemma-4-26b-a4b, deepseek-v3, dots.llm1, glm-5, bitnet, paligemma2) are unaffected and still open.

  4. self-assigned this
    on Jun 17, 2026
  5. inureyes commented on Jun 17, 2026

    @inureyes
    MemberAuthor
  6. inureyes commented on Jun 17, 2026

    @inureyes
    MemberAuthor
  7. inureyes commented on Jun 17, 2026

    @inureyes
    MemberAuthor
  8. inureyes commented on Jun 17, 2026

    @inureyes
    MemberAuthor
  9. added this to the 0.3 milestone on Jun 21, 2026
  10. added and removed on Sep 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:benchmarkBenchmark harness and performance measurement (bench_*.sh, /update-benchmarks)area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)priority:highHigh prioritystatus:doneCompletedtype:bugBug fixes, error corrections, or issue resolutions

Type

No type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions