Skip to content
View nareshns2004's full-sized avatar

Block or report nareshns2004

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
nareshns2004/README.md
A GPU cluster is only as fast as its slowest link and as reliable as its weakest. I make it lose less time to both. Animated: a ring of 8 GPUs loses a link, every rank times out, the job restores onto a spare and resumes.

Website LinkedIn Hugging Face Substack X LeetCode Email

I’m Naresh. I work below the framework layer of AI clusters: kernel networking, RDMA/RoCE, DPDK, SR-IOV, KVM and NCCL, about nine years of it. When a 512-GPU job stalls, the cause is usually one bad link, NIC or GPU, while every rank reports the same timeout. I build the infrastructure that finds that one component and gets training back fast, and that moves bytes between GPUs without wasting the hardware.

Everything here is built in public. A number only appears once it’s measured and linked to its run.

🎯 Two problems I work on

🔵 Reliability at scale

goodput: fault attribution & recovery for multi-node training

Problem. One bad GPU, NIC, cable or switch port stalls a collective, and every rank times out the same way. Stock tooling detects in minutes, attributes nothing, and restarts the whole job.

Approach. Fault injection with a ground-truth ledger, NCCL flight-recorder signatures, DCGM and NIC telemetry joined on the topology to name the culprit rank and component, then the smallest safe recovery: in-process communicator re-init, single-node replacement, peer in-memory checkpoints.

Status. M1 of 6 · design, interfaces and harness public · no results yet

🟠 Efficiency at scale

kvwire: RDMA KV-cache transport for disaggregated inference

Problem. Prefill is compute-bound and decode is memory-bound. Colocated, they interfere: prefills spike TPOT, decodes spike TTFT.

Approach. Split them onto separate GPU pools, and make the part in between fast and safe: a GPUDirect RDMA transport, a Triton KV re-layout kernel (kept only if the data justifies it), and a coordinator that stays correct when the network fails mid-transfer. Plugs into vLLM through its KV-connector interface.

Status. M1 · baselines and hardware ceilings first

▶️ Interactive notes: break things in your browser

Each one is a working simulator or calculator, with its model and assumptions stated. No hand-waved benchmarks.

Post What you can do
💥 One dead link, 512 idle GPUs Cut a link in a live ring all-reduce and watch every rank time out the same way
🔎 Triaging an NCCL timeout Four incident drills: find the culprit from flight-recorder, Xid and NIC evidence
🧊 PFC: the lossless network that can freeze itself Push a toy RoCE fabric into deadlock, then turn on ECN
🛤️ Where NCCL traffic actually goes Rail-optimized vs fat-tree for DP, TP and MoE all-to-all, with and without PXN
📦 The KV-cache bill Size the prefill→decode handoff for real model configs and link speeds
💾 How often should a 16K-GPU job checkpoint? Young/Daly with sliders: which lever buys the most goodput
🧵 The packet path Step one packet through the kernel stack, XDP, DPDK and RDMA

🧱 Where I work in the stack

Where I work in the stack: L5 models & serving (PyTorch, vLLM, Hugging Face, CUDA, Triton kernels, KV cache); L4 collectives & orchestration (NCCL, DP/TP/PP, checkpoint-restart, Ray, Slurm, Kubernetes); L3 transport & fabric (RDMA verbs, RoCE v2, InfiniBand, GPUDirect RDMA, PFC/ECN, rail-optimized fabrics); L2 host datapath (Linux networking, eBPF/XDP, DPDK, SR-IOV, KVM/virtio, perf/ftrace); L1 hardware (NVIDIA GPUs, NVLink, RDMA NICs, PCIe); plus Docker, Prometheus, Grafana, DCGM and C, C++, Python, Go, Bash.
Skill → evidence map (no self-ratings; every row links to the work)
Area Evidence Stage
Collectives & NCCL goodput · distributed-training-framework-nccl · ring all-reduce post design + harness
RDMA / RoCE / PFC nicprof · kvwire · PFC post runnable
Linux kernel performance kptk runnable
Inference serving kvwire · high-performance-llm-inference-engine · KV-cache post early
GPU kernels (CUDA / Triton) custom-cuda-fused-attention-triton early
eBPF / XDP / DPDK kernel-level-ai-traffic-shaper · packet path post · DPDK essay early / write-up
Virtualization (SR-IOV / KVM) virtualization essay write-up
ML for infra ops 6 RCA & failure-prediction models (NCCL, GPU, network, Linux, Kubernetes logs) published

🛠️ Runnable today

  • nicprof: which NIC counter explains which millisecond of training step time. Aligns RDMA/ethtool/SR-IOV counters to each rank’s steps and traces PFC pause back to its origin, so a victim port isn’t blamed. nicprof demo · 69 tests
  • kptk: explains why a Linux workload is slow from kernel evidence: run-queue delay, NUMA locality, THP fallback, and perf_event_open(2) called directly. Standard library only · 43 tests

Also worked with: Java · Kafka · Hadoop · Redis · GraphQL · TensorFlow · Keras · OpenCV · AWS · Terraform · Ansible · Jenkins · Elasticsearch · OpenStack


Open to GPU cluster networking, distributed-training infrastructure and inference-platform roles
USA · Canada · Europe · nareshns2004@gmail.com · portfolio

Pinned Loading

  1. ai-nic-performance-profiler ai-nic-performance-profiler Public

    An observability primitive that closes the attribution gap between NIC hardware counters and distributed training throughput degradation enabling data-driven decisions on fabric topology, RDMA tuni…

    Python

  2. kernel-performance-toolkit kernel-performance-toolkit Public

    A Linux kernel performance analysis toolkit for profiling CPU scheduling, memory behavior, NUMA locality, cache efficiency, page faults and Huge Pages

    Python

  3. distributed-training-framework-nccl distributed-training-framework-nccl Public

    Mini Distributed Training Framework using NCCL

    C++

  4. custom-cuda-fused-attention-triton custom-cuda-fused-attention-triton Public

    Building high-performance GPU kernels from first principles by progressively implementing and optimizing deep learning operators in CUDA and Triton

    Python

  5. disaggregated-inference-engine disaggregated-inference-engine Public

    Python

  6. fault-tolerant-training-orchestrator fault-tolerant-training-orchestrator Public

    A cross-layer reliability system for multi-node LLM training

    Python