I like systems that can prove they work: calibrated models instead of raw scores, LLM output checked against ground truth, pipelines that quarantine bad data instead of hiding it.
M.S. Computer Science, Data Science concentration, UNC Charlotte — May 2026.
Three public repos. Sole author. Each one ships with tests, CI, and a tagged release — or it doesn't ship.
Card-fraud scoring that prices every decision: review a transaction when P(fraud) × amount clears the cost of reviewing it. No fitted threshold, just the expected-cost rule.
Under leak-free chronological evaluation the amount-aware rule saved $84.87 per 10,000 transactions, 95% CI [$20.08, $187.47] — about 13% cheaper than the best single-threshold baseline. It reviews fewer frauds than that baseline and still wins, because it wins on dollars rather than on counts. One finding worth the click: a plain logistic model beat gradient boosting once the evaluation stopped leaking.
Python · scikit-learn · FastAPI · DuckDB · Docker · bootstrap CIs
An evaluation harness, not an extractor. Every US filer publishes its numbers twice — once in prose, once in structured XBRL — so the company's own filing grades the model automatically. No human answer key.
On headline-figure extraction with the financial statements already in the model's context, llama3.1-8B reached 87.8%, 95% CI [81.5, 93.8], across ten large-cap clean us-gaap annual filers. That's the easy end of the task, and the framing travels with the number. A 600-call deterministic grid, 588 scored, all 230 failures hand-labeled into a six-label taxonomy.
Python · XBRL / EDGAR · BM25 · cluster bootstrap · mypy strict
A quality gate for Kafka-compatible streams, measured by what it prevents. Declarative per-topic contracts enforced inline; violations quarantined with self-describing envelopes rather than dropped.
Same faulted fixture replayed through the same bar builder — gates off, 1,076 violated bars of 15,061 got through; gates on, zero, with the output bit-identical to the ground-truth valid projection. Every figure is fixture-scale: a seeded synthetic fixture with a chosen fault mix, never live traffic. 194 tests, 99.4% coverage on core modules, mypy strict.
A stated v0.9, not a 1.0 — the quarantine/replay CLI got cut under a slip valve frozen before the deadline, and the ADR saying so is in the repo. Shipping the roadmap honestly is the point.
Python 3.12 · Kafka / Redpanda · Avro · DuckDB · GitHub Actions
Graduate research assistant, Jan–May 2026. Structural analysis of code refactorings — mutual-information-derived scoring and influence clustering over code property graphs. About 3,500 lines of git-verified contributions. The repo is private faculty research, so the three above are the ones you can actually read.
Instructional assistant, Sep–Dec 2025. Graduate course on software engineering for AI-enabled systems. Office hours, project mentoring across the ML lifecycle, grading.
Earlier. Real-time Apache Kafka + ZooKeeper ingestion and Power BI KPI dashboards in a banking-IT setting, plus a local open-source LLM deployment.
Python · SQL · scikit-learn · pandas · Kafka · Avro · DuckDB · FastAPI · Docker · GitHub Actions · pytest · mypy · Power BI · AWS Glue / S3 · Git
Open to full-time ML Engineer and Data Engineer roles in the U.S.
