I build private AI systems across fintech, edtech, healthcare/pharma, and logistics. I believe in multidisciplinary AI: I apply what works across domains, and I’m always learning.
I’m runtime-, cloud-, and OS-agnostic (vLLM/TGI/TensorRT-LLM/llama.cpp • AWS/GCP/Azure • Linux/macOS/Windows). I pick the right tool for the job, and I believe strong fundamentals make any new stack learnable.
I love open source; I’m not a heavy maintainer, but I advocate for open-source tools and for teams to regain control of their stack in every company I work with.
- Operating private AI LLM inference platforms (on-prem/VPC) end-to-end: runtime selection, KV-cache/VRAM orchestration, quantized weights, batch scheduling; driving >90% GPU utilization
- Building voice AI agents: low-latency streaming ASR/TTS, VAD & barge-in, function/tool use, memory, turn-taking
- Shipping tool-augmented agents: graph/DAG planners, schema-validated JSON outputs, semantic & prompt caching, context compression/model routing (light retrieval only when it helps)
- MLOps/infra: Docker/K8s, CI/CD, tracing/metrics, model registry (MLflow)
- Anime + Video Games
- Reading tech news and arguing on reddit
- Tinkering with new gadgets and stacks






