diff --git a/benchmarks/cuda_gb10_2026-06-17.csv b/benchmarks/cuda_gb10_2026-06-17.csv index ca3cb08d3..0252b81b7 100644 --- a/benchmarks/cuda_gb10_2026-06-17.csv +++ b/benchmarks/cuda_gb10_2026-06-17.csv @@ -3,8 +3,8 @@ apertus-8b-instruct-2509-4bit,./models/apertus-8b-instruct-2509-4bit/,66,31,133. aya-expanse-8b-4bit,./models/aya-expanse-8b-4bit/,8,100,44.74,178.81,2088.40,47.88,2026-06-17,NVIDIA_GB10_122GB,0.3.1,release,100,"Hello, how are you today?" aya-vision-8b,./models/aya-vision-8b/,8,87,64.05,124.91,1864.06,46.67,2026-06-17,NVIDIA_GB10_122GB,0.3.1,release,100,"Hello, how are you today?" baichuan-m1-14b-4bit,./models/baichuan-m1-14b-4bit/,9,7,121.10,74.32,313.06,22.36,2026-06-17,NVIDIA_GB10_122GB,0.3.1,release,100,"Hello, how are you today?" -bitnet-b1.58-2b-4t,./models/bitnet-b1.58-2b-4t/,,,,,,,2026-06-17,NVIDIA_GB10_122GB,0.3.1,release,100,"Hello, how are you today?",FAIL:bench -bitnet-b1.58-2b-4t-4bit,./models/bitnet-b1.58-2b-4t-4bit/,,,,,,,2026-06-17,NVIDIA_GB10_122GB,0.3.1,release,100,"Hello, how are you today?",FAIL:bench +bitnet-b1.58-2b-4t,models/bitnet-b1.58-2b-4t,14,33,27.04,517.84,253.39,130.24,2026-06-17,NVIDIA_GB10_CUDA13.0_122GB,0.3.1,release,100,"Hello, how are you today?" +bitnet-b1.58-2b-4t-4bit,models/bitnet-b1.58-2b-4t-4bit,14,33,26.04,537.54,186.97,176.50,2026-06-17,NVIDIA_GB10_CUDA13.0_122GB,0.3.1,release,100,"Hello, how are you today?" bunny-llama3-8b-4bit,./models/bunny-llama3-8b-4bit/,18,40,41.48,433.97,819.52,48.81,2026-06-17,NVIDIA_GB10_122GB,0.3.1,release,100,"Hello, how are you today?" command-r7b-4bit,./models/command-r7b-4bit/,8,100,62.86,127.27,2092.77,47.78,2026-06-17,NVIDIA_GB10_122GB,0.3.1,release,100,"Hello, how are you today?" deepseek-coder-1.3b-4bit,./models/deepseek-coder-1.3b-4bit/,76,100,16.74,4540.35,1151.31,86.86,2026-06-17,NVIDIA_GB10_122GB,0.3.1,release,100,"Hello, how are you today?" diff --git a/docs/benchmark_results/model_tests.md b/docs/benchmark_results/model_tests.md index 650c5eac3..ee4f04087 100644 --- a/docs/benchmark_results/model_tests.md +++ b/docs/benchmark_results/model_tests.md @@ -84,7 +84,7 @@ For Qwen2.5-0.5B the 4-bit row is the directly comparable cross-hardware figure; | Supported model architectures | 89+ ModelType variants | | Text models tested (M1 Ultra, 2026-06-15) | 136 pass, 2 partial, 4 fail, 9 skip/non-standalone (151 dirs; adds apertus, seed-oss, dots.llm1, granite family, lfm2, plamo-2, falcon-h1, BitNet; diffusiongemma loads via #291) | | Text models tested (M5 Max, 2026-06-15) | 131 pass, 5 partial, 14 fail/skip (0.2.1 full sweep; post-sweep: qwen2.5-vl-3b-4bit fixed by re-download, oversized bf16 hunyuan dropped; neither a code regression) | -| Text models tested (GB10, 2026-06-17) | 133 pass, 11 fail, 3 not-tested/N.A. (glm-5/glm-5.1 weights not downloaded; paligemma2 image-only), 1 OOM-skip (148 total; 0.3.1 with the CUDA fused decode-MoE kernel #319 — 9 MoE models flipped FAIL→pass vs 0.3.0: 5 Qwen MoE + gemma-4-26b-a4b x2 + dots.llm1 + diffusiongemma) | +| Text models tested (GB10, 2026-06-17) | 135 pass, 8 fail, 3 not-tested/N.A. (glm-5/glm-5.1 weights not downloaded; paligemma2 image-only), 2 too-large/capacity (qwen3-next-480b, deepseek-v3) (148 total; 0.3.1 with the CUDA fused decode-MoE kernel #319 + bitnet CUDA kernel #322) | | VLM models tested (GB10, 2026-06-17) | 53 measured image rows (0.3.1) | | VLM models tested (M5 Max, 2026-06-15) | 54 valid VLM rows (0.2.1 full VLM re-sweep; adds qwen3-vl-4b/8b, minicpm-v-4.6-bf16, nemotron-omni, youtu-vl; qwen2.5-vl-3b-4bit restored after re-download) | | VLM models tested (M1 Ultra, 2026-06-15) | 55 measured VLM rows (53 pass + 2 partial) | diff --git a/docs/benchmark_results/model_tests_gb10.md b/docs/benchmark_results/model_tests_gb10.md index 1a9d9b3cc..2b470f425 100644 --- a/docs/benchmark_results/model_tests_gb10.md +++ b/docs/benchmark_results/model_tests_gb10.md @@ -133,7 +133,7 @@ Prefill/Decode are the measured-pass figures from `mlxcel-bench-decode`. Notes r | Model | Status | Prefill (tok/s) | Decode (tok/s) | Notes | |-------|--------|-----------------|----------------|-------| | deepseek-v2-lite-4bit | ✅ | 160.07 | 96.81 | | -| deepseek-v3-4bit | ❌ | - | - | warmup failure (also failed at 0.1.0) | +| deepseek-v3-4bit | ❌ | - | - | too large for GB10 (671B @ 4bit ~350GB > 122GB); present checkpoint incomplete (layers 0-19 of 61) | | dots.llm1.inst-mixed-4-6bit | ✅ | 25.42 | 22.04 | 39 tok | | gpt-oss-120b-4bit | ✅ | 57.75 | 50.48 | 82 tok | | gpt-oss-20b-mxfp4 | ✅ | 126.16 | 77.25 | | @@ -208,8 +208,8 @@ Prefill/Decode are the measured-pass figures from `mlxcel-bench-decode`. Notes r | Model | Status | Prefill (tok/s) | Decode (tok/s) | Notes | |-------|--------|-----------------|----------------|-------| -| bitnet-b1.58-2b-4t | ❌ | - | - | ternary BitNet; fails warmup on CUDA | -| bitnet-b1.58-2b-4t-4bit | ❌ | - | - | ternary BitNet; fails warmup on CUDA | +| bitnet-b1.58-2b-4t | ✅ | 517.84 | 130.24 | CUDA ternary kernel (#322) | +| bitnet-b1.58-2b-4t-4bit | ✅ | 537.54 | 176.50 | CUDA ternary kernel (#322) | ## VLM-capable Models (text-only pass) @@ -334,10 +334,10 @@ Models that accept image input and generated tokens under the `"What is in this | Metric | Count | |--------|-------| | **Total text models attempted** | 148 | -| **Pass (✅)** | 133 | -| **Fail (❌)** | 11 | +| **Pass (✅)** | 135 | +| **Fail (❌)** | 8 | | **Not tested / N.A. (⚪)** | 3 | -| **OOM-skipped (capacity)** | 1 | +| **Too large for GB10 (capacity)** | 2 | | **VLM models measured (image input)** | 53 | ### CUDA fused decode-MoE kernel (#319) @@ -360,12 +360,10 @@ The 0.3.1 line ported the fused decode-MoE kernel to CUDA (#319). Nine MoE model ### Failing / skipped models (by cause) -- **BitNet (ternary; fails CUDA warmup):** `bitnet-b1.58-2b-4t`, `bitnet-b1.58-2b-4t-4bit` -- **Other model-specific failures:** `deepseek-v3-4bit` - **Not tested / not applicable (not a code failure):** `glm-5-4bit`, `glm-5.1-4bit` (weights not downloaded); `paligemma2-3b-6bit` (image-only PaliGemma: 0 text-gen without an image, captions correctly in the VLM image table) - **Not standalone text-gen models:** `docling-layout-heron-mlx-bf16` (document layout), `granite-speech-4.1-2b-nar-mlx` (speech) - **MTP/DFlash drafter checkpoints (need a target; not standalone):** `gemma-4-12b-it-assistant-4bit`, `gemma-4-31b-it-assistant-bf16`, `qwen3.5-27b-dflash`, `qwen3.5-4b-dflash` - **VLM warmup failure under image setup:** `qwen2.5-vl-3b` (bf16), `minicpm-v-4.6-mxfp4` (mxfp4; the bf16 variant passes) -- **OOM-skipped (capacity, weights exceed the memory budget):** `qwen3-next-480b-4bit` +- **Too large for GB10 (capacity, weights exceed the 122 GB budget):** `qwen3-next-480b-4bit`; `deepseek-v3-4bit` (671B @ 4bit ~350GB; the present checkpoint is also an incomplete partial download, layers 0-19 of 61) These remaining failures are tracked in #315.