Most of my work is low-level ML systems work: distributed training, reinforcement learning, agentic systems, inference infrastructure, and GPU optimization.
I've worked on distributed GRPO, SGLang, FSDP2, NCCL, shared-memory transport, TensorRT, vLLM, and multimodal models. In open-source projects, I've built stateful SGLang infrastructure and zero-copy SHM transport reaching 37.8k trajectories/s, implemented wait-free FSDP2 synchronization for distributed rollouts, and built agentic evolution infrastructure with concurrent GRPO/SGLang execution and LoRA synchronization. :chatgpt-content-reference{index="0"}
I've also worked on production AI inference, including TensorRT compilation of Wan2.1 DiT backbones and StreamV2V infrastructure. :chatgpt-content-reference{index="1"}
I tend to work from the systems layer up: profiling bottlenecks, moving data efficiently, keeping state where it matters, and making large-scale training and inference actually run.
Outside ML systems, I've worked on robotics, policy learning, ROS2, and hardware integration. I also have a published paper on Multi-UAV policy learning and formation control.
I'm interested in reinforcement learning, distributed ML, agents, inference systems, and robotics.
I'm also open to select fully remote, contract-based ML research and engineering work.
You can find some of my work in the repositories here.
📫 Personal / engineering: prakarshkaushik369@gmail.com
Turiya / business: sales@turiyahq.com


