Data scientist in Seattle building AI tools, agent systems, and decision products. I care about systems that are useful, inspectable, and honest about where people still need to decide.
leozhao.me · Interested in agent reliability, useful evaluation, and human-in-the-loop product design.
| Project | What it is |
|---|---|
| Loop Agent · demo | A tool-using agent loop with persistent traces, integrity checks, and bounded tool calls, arguments, and results. Runs locally without an API key. |
| Tab Tidy · Chrome Web Store | Chrome extension that proposes tab groups, takes plain-English refinements, and previews every change before applying it. |
| Point2Prompt · install | Bookmarklet that turns a click on any UI element into a structured change brief for Claude Code, Codex, or Cursor. |
| Trackpad Canvas · site | Native macOS diagramming app that draws from raw trackpad touches and snaps sketches into connected architecture diagrams. |
| SplitTaste · demo | Repairs shared-account streaming recommendations with two user questions; reproducible MovieLens 32M evaluation with honest metric gates. |
| Where to Sit · demo | 3D IMAX seat-view simulator for 20 venues across Seattle, NYC, and the Bay Area. |
- Microsoft Agent Framework — surfaced A2A preview consent URLs and added regression coverage.
- Strands Harness SDK — stops retry backoff promptly when a TypeScript run is cancelled.
- OpenMed — adds strict, versioned parsing for agent run summaries.
- DeepEval — normalizes verbose judge verdicts so ambiguous outputs cannot silently become passing scores.
- OpenHarness — prevents disabled tools from leaking into model guidance.

