I’m Naresh. I work below the framework layer of AI clusters: kernel networking, RDMA/RoCE, DPDK, SR-IOV, KVM and NCCL, about nine years of it. When a 512-GPU job stalls, the cause is usually one bad link, NIC or GPU, while every rank reports the same timeout. I build the infrastructure that finds that one component and gets training back fast, and that moves bytes between GPUs without wasting the hardware.
Everything here is built in public. A number only appears once it’s measured and linked to its run.
|
goodput: fault attribution & recovery for multi-node training Problem. One bad GPU, NIC, cable or switch port stalls a collective, and every rank times out the same way. Stock tooling detects in minutes, attributes nothing, and restarts the whole job. Approach. Fault injection with a ground-truth ledger, NCCL flight-recorder signatures, DCGM and NIC telemetry joined on the topology to name the culprit rank and component, then the smallest safe recovery: in-process communicator re-init, single-node replacement, peer in-memory checkpoints. Status. |
kvwire: RDMA KV-cache transport for disaggregated inference Problem. Prefill is compute-bound and decode is memory-bound. Colocated, they interfere: prefills spike TPOT, decodes spike TTFT. Approach. Split them onto separate GPU pools, and make the part in between fast and safe: a GPUDirect RDMA transport, a Triton KV re-layout kernel (kept only if the data justifies it), and a coordinator that stays correct when the network fails mid-transfer. Plugs into vLLM through its KV-connector interface. Status. |
Each one is a working simulator or calculator, with its model and assumptions stated. No hand-waved benchmarks.
| Post | What you can do | |
|---|---|---|
| 💥 | One dead link, 512 idle GPUs | Cut a link in a live ring all-reduce and watch every rank time out the same way |
| 🔎 | Triaging an NCCL timeout | Four incident drills: find the culprit from flight-recorder, Xid and NIC evidence |
| 🧊 | PFC: the lossless network that can freeze itself | Push a toy RoCE fabric into deadlock, then turn on ECN |
| 🛤️ | Where NCCL traffic actually goes | Rail-optimized vs fat-tree for DP, TP and MoE all-to-all, with and without PXN |
| 📦 | The KV-cache bill | Size the prefill→decode handoff for real model configs and link speeds |
| 💾 | How often should a 16K-GPU job checkpoint? | Young/Daly with sliders: which lever buys the most goodput |
| 🧵 | The packet path | Step one packet through the kernel stack, XDP, DPDK and RDMA |
Skill → evidence map (no self-ratings; every row links to the work)
| Area | Evidence | Stage |
|---|---|---|
| Collectives & NCCL | goodput · distributed-training-framework-nccl · ring all-reduce post | design + harness |
| RDMA / RoCE / PFC | nicprof · kvwire · PFC post | runnable |
| Linux kernel performance | kptk | runnable |
| Inference serving | kvwire · high-performance-llm-inference-engine · KV-cache post | early |
| GPU kernels (CUDA / Triton) | custom-cuda-fused-attention-triton | early |
| eBPF / XDP / DPDK | kernel-level-ai-traffic-shaper · packet path post · DPDK essay | early / write-up |
| Virtualization (SR-IOV / KVM) | virtualization essay | write-up |
| ML for infra ops | 6 RCA & failure-prediction models (NCCL, GPU, network, Linux, Kubernetes logs) | published |
- nicprof: which NIC counter explains which millisecond of training step time. Aligns RDMA/ethtool/SR-IOV counters to each rank’s steps and traces PFC pause back to its origin, so a victim port isn’t blamed.
nicprof demo· 69 tests - kptk: explains why a Linux workload is slow from kernel evidence: run-queue delay, NUMA locality, THP fallback, and
perf_event_open(2)called directly. Standard library only · 43 tests
Also worked with: Java · Kafka · Hadoop · Redis · GraphQL · TensorFlow · Keras · OpenCV · AWS · Terraform · Ansible · Jenkins · Elasticsearch · OpenStack
Open to GPU cluster networking, distributed-training infrastructure and inference-platform roles
USA · Canada · Europe · nareshns2004@gmail.com · portfolio

