Skip to content

Add retrieval benchmark harness and A/B config support - #132

Merged
m1rl0k merged 3 commits into
testfrom
benchmarking
Dec 29, 2025
Merged

m1rl0k merged 3 commits into
testfrom
benchmarking

Conversation

@m1rl0k

@m1rl0k m1rl0k commented Dec 29, 2025

Copy link
Copy Markdown
Collaborator

Introduces eval_quality.py for retrieval quality evaluation (Hit@k, MRR), a public benchmark dataset and gold query set, and config file support for A/B testing in run_matrix.py. Adds clone_snapshot.py for reproducible repo cloning, and updates run_matrix.py to support config variants, gold file evaluation, and improved result aggregation.

Introduces eval_quality.py for retrieval quality evaluation (Hit@k, MRR), a public benchmark dataset and gold query set, and config file support for A/B testing in run_matrix.py. Adds clone_snapshot.py for reproducible repo cloning, and updates run_matrix.py to support config variants, gold file evaluation, and improved result aggregation.
@m1rl0k
m1rl0k merged commit 304d0c1 into test Dec 29, 2025
1 check passed
@m1rl0k
m1rl0k deleted the benchmarking branch January 1, 2026 04:21
m1rl0k added a commit that referenced this pull request Mar 1, 2026
Add retrieval benchmark harness and A/B config support
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant