Skip to content
View lazarevtill's full-sized avatar
🏠
Working from home
🏠
Working from home

Organizations

@Lazarev-Cloud

Block or report lazarevtill

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
lazarevtill/README.md

// platform & MLOps engineer

Anatoly (Till) Lazarev

I build the platforms other engineers ship on — ML infrastructure, Kubernetes, GPU, and the security underneath.

Website LinkedIn Email Location


Platform and MLOps engineer, 7+ years. Founding MLOps hire at a fintech serving millions of users, where I designed and built the company's ML platform and still run it as its sole engineer — from the EKS clusters and H100s up to the SDK its users write code against.

The systems on this profile are mine: designed, built and operated end to end, not by a company. Most of my day-to-day code lives on a self-hosted GitLab behind lazarev.cloud; this is the public slice.

What I work on

ML platform, at work. Built from zero: Snowflake and S3 data through notebooks, training, benchmarking and versioning to a served model — Kubeflow, MLflow, KServe, Langfuse and Label Studio on EKS, behind Keycloak SSO and delivered by Argo CD. Five team workspaces run on it.

Two parts I'd point at:

  • Serving with provenance, no tenant credentials. A custom KServe storage container resolves an mlflow:// reference through the tracking server using the team's own token, so no object-storage keys exist in tenant namespaces — with a Kyverno admission policy requiring every InferenceService to reference a registered model version.
  • An SDK, CLI and MCP server. One authenticated entry point for notebooks, training, the registry, serving and status — usable from a laptop over OIDC with no kubeconfig. State-changing operations sit behind an explicit opt-in, because an agent that can redeploy production by accident is a design mistake, not a feature.

GPU capacity runs on HAMi sharing over 8×H100, with selectable VRAM slices and admission-time quota checks that refuse with the actual numbers instead of hanging.

Things I've built

Agentless SSH server monitoring for macOS, iOS, Android, Windows and Linux. One Rust core owns the parsing, scheduling, rate maths and health verdicts; each platform contributes only a view.

A full refresh costs exactly one network round trip however many collectors are enabled — and a test fails if that ever stops being true. Nothing is installed, written or modified on a monitored host; no sample ever reaches disk.

Rust · SwiftUI · Kotlin · GTK4 · WinUI 3 · UniFFI

A project-scoped memory for AI tools, consolidated into dated facts by a local model. Tell it something from any repository; any MCP client recalls it, scoped to the repo it is working in.

Hybrid recall (sqlite-vec + FTS5, Cyrillic-aware) fused by reciprocal rank, facts with validity intervals rather than overwrites, and a labelled probe suite that scores recall@k instead of assuming it works.

Python · SQLite · MCP · llama.cpp · RAG

Measured llama.cpp tuning for AMD Strix Halo — the real memory ceiling (~109 GB, not the 96 GB the BIOS implies), the flags that matter, and five "obvious" optimisations that measurement killed.

Shipped with the eval harness and a catalogue of fourteen harness bugs, each with the believable wrong number it produced. When the suites turned out to be measuring the tasks instead of the models, the numbers were withdrawn rather than published.

PowerShell · Python · Vulkan · benchmarking

My own production platform: a 9-node Proxmox cluster with NVIDIA GPU and Ryzen AI NPU passthrough for local ML workloads, fully infrastructure-as-code across 11 OpenTofu providers.

Vault internal PKI and SSO everywhere, a self-hosted CI/CD supply chain (GitLab, Harbor, Nexus, Renovate), Prometheus/Grafana observability, 3-2-1 backups.

Proxmox · OpenTofu · Vault · GitLab CI · Harbor

Also: fingerprint-manager, a Qt desktop front end for fprintd.

Upstream contributions

Fixes I found while running the ML platform I built in production, sent back to the projects it is built on.

Project Change Status
MLflow Eagerly load tags, params and metrics in search_logged_models — one page cost 3N+1 queries (~300 round trips at the default page size); now a constant number ✅ merged
MLflow Stop a metric filter from shrinking search_logged_models pages — the filter join duplicated rows before LIMIT, so filtered pages came back short ✅ merged
MLflow Let search_logged_models callers skip metric values — measured on a real experiment: 900 ms / 7.15 MB → 12.9 ms / 0.01 MB for the runs table 🔄 in review
mlflow-oidc-auth Serve plugin endpoints under /ajax-api as well as /api — the MLflow UI got 404s from the auth plugin ✅ merged
mlflow-oidc-auth Skip the protobuf round-trip when a logged-model page needs no filtering ✅ merged
Kubeflow Pipelines Make the compiled-template Istio sidecar default configurable — on STRICT mTLS meshes no pipeline could run 🔄 in review

Background

Seven years of secure, scalable platforms across bare-metal Linux and AWS. Before the founding MLOps role I was DevOps Manager at the same fintech, standardising Kubernetes across multi-region clusters for 6+ teams and cutting allocated CPU and memory by ~10%. Earlier: a 50%+ cut in company-wide AWS spend and 60% faster deployments at a Dubai real-estate group, and observability handling 150,000 metrics per second at a US e-commerce company.

Toolbox

Kubernetes (EKS & bare-metal) Kubeflow MLflow KServe Istio Kyverno HAMi / NVIDIA H100 AWS Snowflake OpenTofu / Terraform Helm Argo CD GitLab CI Proxmox HashiCorp Vault (PKI/OIDC) Keycloak CrowdSec Prometheus / VictoriaMetrics / Grafana OpenTelemetry Python Go Rust Bash Linux Cisco networking


lazarev.cloud · LinkedIn · till@lazarev.cloud

Just a man with servers and ideas.

Pinned Loading

  1. Morgan Morgan Public

    Self-hosted, self-learning, provider-agnostic personal agent kernel. Local-first memory in one SQLite database, project-scoped, reachable from any tool over MCP or a CLI.

    Python 6

  2. ServerGlass ServerGlass Public

    Agentless SSH server monitoring for macOS, iOS, Android and Linux. One Rust core does the parsing, scheduling and health verdicts; each platform contributes only a view. One network round trip per …

    Rust 1

  3. Commitsmith Commitsmith Public

    Conventional Commit messages from your git diff, in VS Code and Cursor - using any OpenAI-compatible API, Ollama or llama.cpp. Local-first, no telemetry, no code indexing.

    TypeScript 1

  4. strix-halo-llm strix-halo-llm Public

    Measured llama.cpp Vulkan tuning for AMD Strix Halo (Ryzen AI MAX+ 395 / gfx1151), plus an eval harness with a catalogue of 13 bugs that each produced a believable wrong number.

    PowerShell 1 1

  5. heimdall heimdall Public

    Deterministic three-tier log and metric observability for a homelab: signals evaluated against an IaC-generated inventory of expected state, delivered via Alertmanager and YouTrack. A local LLM rea…

    Go 1

  6. fingerprint-manager fingerprint-manager Public

    Add, test and remove fingerprints — a desktop front end for fprintd

    Python