Skip to content

docs(bench): reclassify GB10 0.3.1 results; minicpm-v-mxfp4 passes, drop qwen2.5-vl-3b - #335

Merged
inureyes merged 1 commit into
mainfrom
docs/gb10-bench-reclassify-0.3.1
Jun 17, 2026
Merged

inureyes merged 1 commit into
mainfrom
docs/gb10-bench-reclassify-0.3.1

Conversation

@inureyes

Copy link
Copy Markdown
Member

Follows the loader fix (#334). Brings the GB10 0.3.1 benchmark CSVs and docs in line with the now-resolved results.

Changes

  • minicpm-v-4.6-mxfp4 → pass (was a warmup abort, fixed in fix(vlm): load non-affine VLM weights with the right quant mode and group_size #334). Text 347.59 prefill / 142.14 decode tok/s; image input 625.47 / 138.70 (48 tok, early-EOS). Added to the image-input table; flipped to ✅ in the text table.
  • qwen2.5-vl-3b → removed. Broken duplicate of qwen2.5-vl-3b-4bit (the model dir was deleted upstream); the working 4bit variant stays. Both CSVs go 148 → 147 rows.
  • Six by-design non-failures reclassified ❌ → ⚪ (N.A., not a code failure): docling-layout-heron-mlx-bf16 and granite-speech-4.1-2b-nar-mlx (not standalone text generators), and the four MTP/DFlash drafters gemma-4-12b-it-assistant-4bit, gemma-4-31b-it-assistant-bf16, qwen3.5-27b-dflash, qwen3.5-4b-dflash (need a target model). The raw CSV FAIL:bench markers are unchanged (they genuinely emit no text-gen output); ⚪ is a doc-layer classification, matching how paligemma2 is already handled.

GB10 0.3.1 summary (now)

Metric Count
Total text models attempted 147
Pass (✅) 136
Fail (❌, code failure) 0
Not tested / N.A. (⚪) 9
Too large for GB10 (capacity, OOM ❌) 2
VLM models measured (image input) 54

136 + 0 + 9 + 2 = 147. Verified against the table rows (text-family 136 ✅ / 2 ❌ / 9 ⚪; image table 54 ✅). Legend and the by-cause section rewritten to match; index counts and the benchmark-CSV table updated in step.

…rop qwen2.5-vl-3b

Follows the loader fix (#334). Updates the GB10 0.3.1 benchmark CSVs and docs:

- minicpm-v-4.6-mxfp4: now passes after #334. Text 347.59 prefill /
  142.14 decode tok/s, image input 625.47 / 138.70 (48 tok). Added to the
  image-input table and flipped to a pass in the text table.
- qwen2.5-vl-3b: removed. It was a broken duplicate of qwen2.5-vl-3b-4bit
  (the model dir was deleted); the working 4bit variant stays. Both CSVs
  go 148 -> 147 rows.
- Reclassified six by-design non-failures from fail (red x) to N.A.
  (white circle): docling-layout-heron and granite-speech (not standalone
  text generators) and the four MTP/DFlash drafters (gemma-4-12b/31b
  assistant, qwen3.5-27b/4b-dflash) that need a target model. The raw CSV
  FAIL:bench markers stay (they genuinely produce no text-gen output); the
  N.A. tag is a doc-layer classification.

GB10 0.3.1 summary is now 147 total: 136 pass, 0 code failures, 9 N.A.,
2 too-large (qwen3-next-480b, deepseek-v3, shown red x, OOM), 54 VLM image
rows. Legend and the by-cause section were rewritten to match; index
counts and the benchmark-CSV table updated in step.
@inureyes inureyes added status:review Under review type:docs Documentation improvements or additions labels Jun 17, 2026
@inureyes
inureyes merged commit 918a6ee into main Jun 17, 2026
5 checks passed
@inureyes
inureyes deleted the docs/gb10-bench-reclassify-0.3.1 branch June 17, 2026 05:36
@inureyes inureyes added this to the 0.3 milestone Jun 21, 2026
@inureyes inureyes self-assigned this Aug 31, 2026
@inureyes inureyes added status:done Completed and removed status:review Under review labels Sep 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

status:done Completed type:docs Documentation improvements or additions

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant