Repository navigation
feat(eval): A/B eval runner (Claude Code with vs without Astrograph) - #12
Merged
Merged
Conversation
`bun run eval [filter...]` runs each task twice with `claude -p`, once with the Astrograph MCP server and skill and once without, and records tool calls, tokens, cost and a keyword score per run. - Repositories are pinned to a commit and checked out fresh under ~/.cache/astrograph-eval, so working copies are never touched. - eval/repos/*.json (public) is committed; eval/local/*.json (private repos and their tasks) is gitignored and skipped when the path is missing. - Both arms run with a clean environment, read-only tools, no user settings or skills, and only the MCP servers the eval passes. Part of #3 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #3
What
bun run eval [filter...]answers each task twice withclaude -p— once with Astrograph (MCP server + skill, exactly whatastrograph installsets up), once without — and records tool calls, Astrograph calls, tokens, cost and a keyword score per run. Transcripts andruns.jsongo toeval/results/<timestamp>/(gitignored).scripts/repos.ts: reference repositories pinned to a commit and checked out fresh under~/.cache/astrograph-eval(git fetch --depth 1 <url|path> <sha>). Working copies are never touched; uncommitted changes in them are ignored..astrograph/and the eval's skill are added to the clone's.git/info/excludeso grep/rg don't see them. Run bench in CI on pinned public repos #8 will reuse this for the CI bench.eval/repos/*.json(committed): umami, trpc, t3code, koel, fastify — 3 tasks each (trace / callers / impact), answers verified against the source at the pinned commit.eval/local/*.json(gitignored): private repositories and their tasks, same shape withpathinstead ofurl. A missing path is reported (skip <name>: not found at …) and skipped, so anyone can run the public set with no setup.Isolation (why the comparison is fair)
Found while smoke-testing: the first version leaked the calling session's
CLAUDE_CODE_*/CLAUDE_EFFORTenv and user skills (including an installed astrograph skill) into both arms, and--allowedToolsalone did not stopawkfrom running.Both arms now run with:
--setting-sources project(no user settings/skills),--strict-mcp-config(only the MCP server the eval passes);--tools Read,Grep,Glob,Bash,ToolSearch,Skill+--permission-mode dontAsk+ an allowlist of read-only Bash commands.The arm with Astrograph additionally gets
--mcp-config(server run from this checkout, so the eval measures the branch) and the skill copied into the clone as a project skill.Verified
bun test,bun run typecheck,bun run checkgreen.astrograph: connected+ skill in the arm with,permissionMode: dontAsk.🤖 Generated with Claude Code