Peak local-inference fork: Qwen3.8-27B @ 196K + native MTP on single 16GB GPU (RTX 5070 Ti). Gated benchmark suite (needle/toolcall/coherence/speed). den-legacy branch holds prior engine.
-
Updated
Oct 11, 2026 - C++
Peak local-inference fork: Qwen3.8-27B @ 196K + native MTP on single 16GB GPU (RTX 5070 Ti). Gated benchmark suite (needle/toolcall/coherence/speed). den-legacy branch holds prior engine.
NCCL over a switchless 4-node DGX Spark / GB10 ring using both PCIe halves of every QSFP cable: ~193 Gb/s per cable instead of ~112, up to +34% vLLM prefill. One patch on top of switchless-nccl.
To associate your repository with the nvfp4-native topic, visit your repo's landing page and select "manage topics."