Repository navigation
docs(bench): GB10 0.3.1 benchmark results - #320
Merged
Merged
Conversation
First GB10 (DGX Spark, CUDA 13.0 / SM 12.1) benchmark on the merged CUDA fused decode-MoE kernel (#319): 148 text models (133 pass / 14 fail / 1 OOM-skip), 53 VLM image rows. Nine fused-MoE models that aborted at 0.3.0 on the Metal-only kernel now run on CUDA, and qwen3-moe-30b (89.84) edges past M1 Ultra (83.75). Regenerate model_tests_gb10.md and refresh the cross-hardware index for 0.3.1. The Cargo.toml bump to 0.3.1 lands at release, so the CSVs are relabeled to 0.3.1 to match the line they measure.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
First GB10 (DGX Spark, CUDA 13.0 / SM 12.1) benchmark on the merged CUDA fused decode-MoE kernel (#319).
Text: 148 models — 133 pass / 14 fail / 1 OOM-skip (vs 0.3.0: 124 / 23 / 1). Nine fused-MoE models that aborted at 0.3.0 on the Metal-only kernel now run on CUDA:
qwen3-moe-30b(89.84) now edges past M1 Ultra (83.75). VLM: 53 image rows (vs 49).Regenerates
model_tests_gb10.mdand refreshes the cross-hardware index for 0.3.1. The Cargo.toml bump lands at release, so the CSVs are relabeled to 0.3.1 to match the line they measure. Remaining failures (deepseek-v3, glm-5, bitnet, paligemma, drafters, VLM variants) are tracked in #315.