Open-source AI engineer in Ho Chi Minh City, working on diffusion/video generation, 4-bit diffusion inference, LLM/VLM quantization, model serving, and practical ML infrastructure.
GitHub | Hugging Face | Website
Diffusion · Video generation · SVDQuant / 4-bit inference · LLM/VLM quantization · Model serving · Open-source ML
| Project | ✓ Merged PRs | Representative merged work |
|---|---|---|
| huggingface/diffusers ★ 34,623 / forks 7,356 |
24 | Nunchaku Lite quantization backend, LTX2 distilled checkpoint support, framewise LTX Video VAE encoding/decoding |
| huggingface/transformers ★ 166,749 / forks 34,699 |
8 | PoolFormer fast image processor, BridgeTower fast image processor |
| huggingface/blog | 1 | Bringing Nunchaku 4-bit Diffusion Inference to Diffusers, co-authored with the Diffusers team |
| sgl-project/sglang ★ 36,508 / forks 9,189 |
1 | Z-Image text encoder config fix |
| d2l-ai/d2l-vi ★ 663 / forks 254 |
108 | Vietnamese ML education translation/revision |
| mlbvn/ml-yearning-vi ★ 1,234 / forks 395 |
23 | Vietnamese ML education translation/revision |
| d2l-ai/d2l-en ★ 29,724 / forks 5,140 |
7 | Documentation fixes |
Selected upstream work includes the Nunchaku Lite 4-bit quantization backend in Diffusers, LTX2 distilled checkpoint support, LTX Video VAE framewise encoding/decoding, Diffusers pipeline/test improvements, a Z-Image fix in SGLang, and Vietnamese ML education work across Dive into Deep Learning and Machine Learning Yearning.
| Project | What it does |
|---|---|
| nunchaku-lite | Lean runtime for SVDQuant W4A4 diffusion models in standard Diffusers pipelines, with native CUDA kernels published as nunchaku-lite-kernels on the Hub and torch.compile support |
| diffuse-compressor | Model-agnostic SVDQuant toolkit that calibrates and quantizes diffusion transformers to INT4 / NVFP4 and exports Nunchaku-compatible checkpoints |
| lite-infer | Hub organization with 30 prequantized Nunchaku Lite checkpoints: FLUX.1 / FLUX.2 Klein, Qwen-Image and Qwen-Image-Edit, Z-Image-Turbo, ERNIE-Image-Turbo, Krea-2-Turbo, and LTX-2.3 Distilled |
| diffuser_layerdiffuse | Transparent image generation with Diffusers (LayerDiffuse) |
| Project | Where it is used | Adopting project scale |
|---|---|---|
| rootonchair/nunchaku-lite | Diffusers: official quantization backend with a dedicated docs page SD.Next: built-in Nunchaku-Lite inference engine that installs the package and ships 11 lite-infer checkpoints in its model reference |
★ 34,623 / forks 7,356 ★ 7,346 / forks 585 |
| rootonchair/LTX-2-19b-distilled | Listed in vLLM Omni's supported-model table for LTX2DistilledOneStagePipeline and LTX2DistilledTwoStagePipeline |
★ 7,101 / forks 1,802 |
| rootonchair/diffuser_layerdiffuse | SD.Next includes a LayerDiffuse: Transparent Image extension and links back to this project in the extension UI |
★ 7,346 / forks 585 |
- Video: Diffusers integrations, LTX video support, LTX-2.3 Distilled Diffusers conversions, transparent image generation, and model execution workflows.
- 4-bit diffusion: SVDQuant INT4 / NVFP4 pipelines for image and video transformers, including MiniMax-H3, Ideogram v4, and ERNIE-Image-Turbo.
- VLM quantization: GGUF and AWQ releases for image-text models such as Lens-Turbo, Vintern-3B, Vintern-1B, EraX-VL-7B, and InternVL2.5-4B.
- Serving: quantized/GGUF model artifacts, runtime tooling, llama.cpp-based OCR serving, and diffusion model compression experiments.
- Tooling: practical patches across Hugging Face Diffusers/Transformers and related open-source ML projects.
I am mostly interested in making generative models easier to run, adapt, compress, and serve in real systems.






