Skip to content
#

dflash2

Here are 23 public repositories matching this topic...

Serving Qwen3.8-27B-FP8 on a single DGX Spark (GB10): 7.88 to 58.5 tok/s single-stream from decode strategy alone, weights untouched. Speculative decoding and prefix caching benchmarked, plus DFlash 2 — the only Qwen3.8-27B build that can serve it under vLLM.

  • Updated Sep 29, 2026
  • Python

TensorFold v0.6.0 GLM-5.3-Flash engine generalized from 2 to 4 DGX Sparks (TP4) over a switchless RoCE fiber ring - uncensored EXL3 checkpoint, overlay + tooling. Upstreams: ashhart/TensorFold, MiaAI-Lab 2x-Spark stack, alexellis/switchless-nccl.

  • Updated Oct 3, 2026
  • Python

Add this topic to your repo

To associate your repository with the dflash2 topic, visit your repo's landing page and select "manage topics."

Learn more