- [2026-06-10] DefTruth, Butterfingrz (2026). FFPA: Efficient Flash Prefill Attention for Large Head Dimensions via Split-D. Zenodo, 2026. 🎉🎉🎉
xlite-dev
Develop ML/AI toolkits and ML/AI/CUDA Learning resources.
Pinned Loading
Repositories
Showing 10 of 73 repositories
- LeetCUDA Public
📚LeetCUDA: Modern CUDA Learn Notes with PyTorch for Beginners🐑, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.🎉
- ffpa-attn Public
🤖FFPA: Extends FA-2/3 via Split-D for large headdims, 1.5x~6×↑🎉 vs SDPA, up to 513~535 TFLOPS🎉 on NVIDIA H200.
- .github Public
- sglang Public Forked from sgl-project/sglang
SGLang is a fast serving framework for large language models and vision language models.
- cudnn-frontend Public Forked from NVIDIA/cudnn-frontend
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
- cutlass Public Forked from NVIDIA/cutlass
CUDA Templates and Python DSLs for High-Performance Linear Algebra
- TurboDiffusion Public Forked from thu-ml/TurboDiffusion
TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
- flash-attention Public Forked from Dao-AILab/flash-attention
Fast and memory-efficient exact attention
Top languages
Loading…
Most used topics
Loading…
