Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions benchmarks/cuda_gb10_2026-06-17.csv
Original file line number Diff line number Diff line change
Expand Up @@ -3,8 +3,8 @@ apertus-8b-instruct-2509-4bit,./models/apertus-8b-instruct-2509-4bit/,66,31,133.
aya-expanse-8b-4bit,./models/aya-expanse-8b-4bit/,8,100,44.74,178.81,2088.40,47.88,2026-06-17,NVIDIA_GB10_122GB,0.3.1,release,100,"Hello, how are you today?"
aya-vision-8b,./models/aya-vision-8b/,8,87,64.05,124.91,1864.06,46.67,2026-06-17,NVIDIA_GB10_122GB,0.3.1,release,100,"Hello, how are you today?"
baichuan-m1-14b-4bit,./models/baichuan-m1-14b-4bit/,9,7,121.10,74.32,313.06,22.36,2026-06-17,NVIDIA_GB10_122GB,0.3.1,release,100,"Hello, how are you today?"
bitnet-b1.58-2b-4t,./models/bitnet-b1.58-2b-4t/,,,,,,,2026-06-17,NVIDIA_GB10_122GB,0.3.1,release,100,"Hello, how are you today?",FAIL:bench
bitnet-b1.58-2b-4t-4bit,./models/bitnet-b1.58-2b-4t-4bit/,,,,,,,2026-06-17,NVIDIA_GB10_122GB,0.3.1,release,100,"Hello, how are you today?",FAIL:bench
bitnet-b1.58-2b-4t,models/bitnet-b1.58-2b-4t,14,33,27.04,517.84,253.39,130.24,2026-06-17,NVIDIA_GB10_CUDA13.0_122GB,0.3.1,release,100,"Hello, how are you today?"
bitnet-b1.58-2b-4t-4bit,models/bitnet-b1.58-2b-4t-4bit,14,33,26.04,537.54,186.97,176.50,2026-06-17,NVIDIA_GB10_CUDA13.0_122GB,0.3.1,release,100,"Hello, how are you today?"
bunny-llama3-8b-4bit,./models/bunny-llama3-8b-4bit/,18,40,41.48,433.97,819.52,48.81,2026-06-17,NVIDIA_GB10_122GB,0.3.1,release,100,"Hello, how are you today?"
command-r7b-4bit,./models/command-r7b-4bit/,8,100,62.86,127.27,2092.77,47.78,2026-06-17,NVIDIA_GB10_122GB,0.3.1,release,100,"Hello, how are you today?"
deepseek-coder-1.3b-4bit,./models/deepseek-coder-1.3b-4bit/,76,100,16.74,4540.35,1151.31,86.86,2026-06-17,NVIDIA_GB10_122GB,0.3.1,release,100,"Hello, how are you today?"
Expand Down
2 changes: 1 addition & 1 deletion docs/benchmark_results/model_tests.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,7 @@ For Qwen2.5-0.5B the 4-bit row is the directly comparable cross-hardware figure;
| Supported model architectures | 89+ ModelType variants |
| Text models tested (M1 Ultra, 2026-06-15) | 136 pass, 2 partial, 4 fail, 9 skip/non-standalone (151 dirs; adds apertus, seed-oss, dots.llm1, granite family, lfm2, plamo-2, falcon-h1, BitNet; diffusiongemma loads via #291) |
| Text models tested (M5 Max, 2026-06-15) | 131 pass, 5 partial, 14 fail/skip (0.2.1 full sweep; post-sweep: qwen2.5-vl-3b-4bit fixed by re-download, oversized bf16 hunyuan dropped; neither a code regression) |
| Text models tested (GB10, 2026-06-17) | 133 pass, 11 fail, 3 not-tested/N.A. (glm-5/glm-5.1 weights not downloaded; paligemma2 image-only), 1 OOM-skip (148 total; 0.3.1 with the CUDA fused decode-MoE kernel #319 — 9 MoE models flipped FAIL→pass vs 0.3.0: 5 Qwen MoE + gemma-4-26b-a4b x2 + dots.llm1 + diffusiongemma) |
| Text models tested (GB10, 2026-06-17) | 135 pass, 8 fail, 3 not-tested/N.A. (glm-5/glm-5.1 weights not downloaded; paligemma2 image-only), 2 too-large/capacity (qwen3-next-480b, deepseek-v3) (148 total; 0.3.1 with the CUDA fused decode-MoE kernel #319 + bitnet CUDA kernel #322) |
| VLM models tested (GB10, 2026-06-17) | 53 measured image rows (0.3.1) |
| VLM models tested (M5 Max, 2026-06-15) | 54 valid VLM rows (0.2.1 full VLM re-sweep; adds qwen3-vl-4b/8b, minicpm-v-4.6-bf16, nemotron-omni, youtu-vl; qwen2.5-vl-3b-4bit restored after re-download) |
| VLM models tested (M1 Ultra, 2026-06-15) | 55 measured VLM rows (53 pass + 2 partial) |
Expand Down
16 changes: 7 additions & 9 deletions docs/benchmark_results/model_tests_gb10.md
Original file line number Diff line number Diff line change
Expand Up @@ -133,7 +133,7 @@ Prefill/Decode are the measured-pass figures from `mlxcel-bench-decode`. Notes r
| Model | Status | Prefill (tok/s) | Decode (tok/s) | Notes |
|-------|--------|-----------------|----------------|-------|
| deepseek-v2-lite-4bit | ✅ | 160.07 | 96.81 | |
| deepseek-v3-4bit | ❌ | - | - | warmup failure (also failed at 0.1.0) |
| deepseek-v3-4bit | ❌ | - | - | too large for GB10 (671B @ 4bit ~350GB > 122GB); present checkpoint incomplete (layers 0-19 of 61) |
| dots.llm1.inst-mixed-4-6bit | ✅ | 25.42 | 22.04 | 39 tok |
| gpt-oss-120b-4bit | ✅ | 57.75 | 50.48 | 82 tok |
| gpt-oss-20b-mxfp4 | ✅ | 126.16 | 77.25 | |
Expand Down Expand Up @@ -208,8 +208,8 @@ Prefill/Decode are the measured-pass figures from `mlxcel-bench-decode`. Notes r

| Model | Status | Prefill (tok/s) | Decode (tok/s) | Notes |
|-------|--------|-----------------|----------------|-------|
| bitnet-b1.58-2b-4t | ❌ | - | - | ternary BitNet; fails warmup on CUDA |
| bitnet-b1.58-2b-4t-4bit | ❌ | - | - | ternary BitNet; fails warmup on CUDA |
| bitnet-b1.58-2b-4t | ✅ | 517.84 | 130.24 | CUDA ternary kernel (#322) |
| bitnet-b1.58-2b-4t-4bit | ✅ | 537.54 | 176.50 | CUDA ternary kernel (#322) |

## VLM-capable Models (text-only pass)

Expand Down Expand Up @@ -334,10 +334,10 @@ Models that accept image input and generated tokens under the `"What is in this
| Metric | Count |
|--------|-------|
| **Total text models attempted** | 148 |
| **Pass (✅)** | 133 |
| **Fail (❌)** | 11 |
| **Pass (✅)** | 135 |
| **Fail (❌)** | 8 |
| **Not tested / N.A. (⚪)** | 3 |
| **OOM-skipped (capacity)** | 1 |
| **Too large for GB10 (capacity)** | 2 |
| **VLM models measured (image input)** | 53 |

### CUDA fused decode-MoE kernel (#319)
Expand All @@ -360,12 +360,10 @@ The 0.3.1 line ported the fused decode-MoE kernel to CUDA (#319). Nine MoE model

### Failing / skipped models (by cause)

- **BitNet (ternary; fails CUDA warmup):** `bitnet-b1.58-2b-4t`, `bitnet-b1.58-2b-4t-4bit`
- **Other model-specific failures:** `deepseek-v3-4bit`
- **Not tested / not applicable (not a code failure):** `glm-5-4bit`, `glm-5.1-4bit` (weights not downloaded); `paligemma2-3b-6bit` (image-only PaliGemma: 0 text-gen without an image, captions correctly in the VLM image table)
- **Not standalone text-gen models:** `docling-layout-heron-mlx-bf16` (document layout), `granite-speech-4.1-2b-nar-mlx` (speech)
- **MTP/DFlash drafter checkpoints (need a target; not standalone):** `gemma-4-12b-it-assistant-4bit`, `gemma-4-31b-it-assistant-bf16`, `qwen3.5-27b-dflash`, `qwen3.5-4b-dflash`
- **VLM warmup failure under image setup:** `qwen2.5-vl-3b` (bf16), `minicpm-v-4.6-mxfp4` (mxfp4; the bf16 variant passes)
- **OOM-skipped (capacity, weights exceed the memory budget):** `qwen3-next-480b-4bit`
- **Too large for GB10 (capacity, weights exceed the 122 GB budget):** `qwen3-next-480b-4bit`; `deepseek-v3-4bit` (671B @ 4bit ~350GB; the present checkpoint is also an incomplete partial download, layers 0-19 of 61)

These remaining failures are tracked in #315.