Skip to content

feat(eval): A/B eval runner (Claude Code with vs without Astrograph) - #12

Merged
Chofito merged 1 commit into
mainfrom
feat/eval-runner
Oct 6, 2026
Merged

Chofito merged 1 commit into
mainfrom
feat/eval-runner

Conversation

@Chofito

@Chofito Chofito commented Oct 6, 2026

Copy link
Copy Markdown
Owner

Part of #3

What

bun run eval [filter...] answers each task twice with claude -p — once with Astrograph (MCP server + skill, exactly what astrograph install sets up), once without — and records tool calls, Astrograph calls, tokens, cost and a keyword score per run. Transcripts and runs.json go to eval/results/<timestamp>/ (gitignored).

  • scripts/repos.ts: reference repositories pinned to a commit and checked out fresh under ~/.cache/astrograph-eval (git fetch --depth 1 <url|path> <sha>). Working copies are never touched; uncommitted changes in them are ignored. .astrograph/ and the eval's skill are added to the clone's .git/info/exclude so grep/rg don't see them. Run bench in CI on pinned public repos #8 will reuse this for the CI bench.
  • eval/repos/*.json (committed): umami, trpc, t3code, koel, fastify — 3 tasks each (trace / callers / impact), answers verified against the source at the pinned commit.
  • eval/local/*.json (gitignored): private repositories and their tasks, same shape with path instead of url. A missing path is reported (skip <name>: not found at …) and skipped, so anyone can run the public set with no setup.

Isolation (why the comparison is fair)

Found while smoke-testing: the first version leaked the calling session's CLAUDE_CODE_*/CLAUDE_EFFORT env and user skills (including an installed astrograph skill) into both arms, and --allowedTools alone did not stop awk from running.
Both arms now run with:

  • a clean env (PATH, HOME, … only), --setting-sources project (no user settings/skills), --strict-mcp-config (only the MCP server the eval passes);
  • --tools Read,Grep,Glob,Bash,ToolSearch,Skill + --permission-mode dontAsk + an allowlist of read-only Bash commands.

The arm with Astrograph additionally gets --mcp-config (server run from this checkout, so the eval measures the branch) and the skill copied into the clone as a project skill.

Verified

  • bun test, bun run typecheck, bun run check green.
  • Smoke runs on fastify: init events confirm 0 MCP servers / no astrograph skill in the arm without, astrograph: connected + skill in the arm with, permissionMode: dontAsk.
  • All 11 repositories (5 public + 6 private) materialize and index; the full run is posted to A/B eval: agent with vs without Astrograph #3.

🤖 Generated with Claude Code

`bun run eval [filter...]` runs each task twice with `claude -p`, once with the
Astrograph MCP server and skill and once without, and records tool calls,
tokens, cost and a keyword score per run.

- Repositories are pinned to a commit and checked out fresh under
  ~/.cache/astrograph-eval, so working copies are never touched.
- eval/repos/*.json (public) is committed; eval/local/*.json (private repos
  and their tasks) is gitignored and skipped when the path is missing.
- Both arms run with a clean environment, read-only tools, no user settings
  or skills, and only the MCP servers the eval passes.

Part of #3

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@Chofito
Chofito merged commit b116122 into main Oct 6, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant