You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
Complete GB10 (DGX Spark, sm_121) block-FP8 Triton kernel configs for DeepSeek-V4-Flash-0731 at TP=2, including the DSpark main_proj shape missing from every earlier run
Speculative-draft training pipeline (DFLASH/DSPARK/DFLASH2), orchestrating the vllm-project/speculators engine end to end, with preset configs for Qwen-family models