Repository navigation
perf(bitnet): CUDA port of the BitLinear ternary matmul kernel - #322
Merged
Merged
Conversation
The BitLinear ternary matmul (BitNet b1.58) was a Metal-only mx.fast.metal_kernel, so on the CUDA backend it aborted at warmup with "[metal_kernel] No Metal back-end" for both bitnet-b1.58-2b-4t and the 4bit variant. Port it to mx.fast.cuda_kernel: one warp per (batch, out/4) row group, the simd_sum reduction over in_features becomes __shfl_down_sync, and the 2-bit-packed ternary unpacking (4 output rows per byte) is unchanged. bitlinear_matmul selects the cuda_kernel port when metal::is_available() is false. Verified on GB10 (DGX Spark, CUDA 13.0): both bitnet variants now produce coherent output (previously an abort).
inureyes
added a commit
that referenced
this pull request
Jun 17, 2026
…large (#324) bitnet re-benched on the fixed binary (passes, merged into CSV); deepseek-v3 reclassified as too-large/capacity (671B ~350GB, incomplete download). Counts 135/8/3/2.
inureyes
added a commit
that referenced
this pull request
Jun 17, 2026
The two bitnet rows were measured separately when their CUDA ternary kernel landed (#322), so they carried a different hardware label (NVIDIA_GB10_CUDA13.0_122GB vs NVIDIA_GB10_122GB) and an unprefixed model_path (models/... vs ./models/.../) from the other 145 rows. Same machine and run, just label drift. Normalized both columns so all 147 rows are uniform. Measured numbers are unchanged.
5 of 11 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The BitLinear ternary matmul (BitNet b1.58) was a Metal-only
mx.fast.metal_kernel, so on the CUDA backend bothbitnet-b1.58-2b-4tand the 4bit variant aborted at warmup with[metal_kernel] No Metal back-end.Port the kernel to
mx.fast.cuda_kernel(the CUDA analogue ofmetal_kernel), the same pattern as #319 (the fused decode-MoE CUDA port):simd_sumreduction overin_featuresbecomes a__shfl_down_syncwarp reduction.bitlinear_matmulselects the cuda_kernel port whenmetal::is_available()is false;no_cuda.cpp/no_metal.cppstub the unused side, so both link on either backend.Validation (GB10 / DGX Spark, CUDA 13.0)
Both bitnet variants now produce coherent output (previously an abort):
Only
src/lib/mlxcel-core/cpp/mlx_cxx_kernels.cppchanges.