Research benchmark for evidence-grounded OS-agent collaboration, continuous state diagnosis, scoped memory reuse, and stale-state rejection.
-
Updated
Jul 29, 2026 - JavaScript
Research benchmark for evidence-grounded OS-agent collaboration, continuous state diagnosis, scoped memory reuse, and stale-state rejection.
Agent-LLM Benchmark Model Results
AI business-logic audit benchmark for Codex-style coding agents
用同一套提示词评测多模型 Agent 从零实现完整可玩游戏的能力:固定技术栈、固定评分标准、按模型归档对比。| A standardized prompt for benchmarking LLM agents on building a complete game from scratch — fixed tech stack, fixed rubric, per-model results.
A long-horizon benchmark for agents that must keep a civilization alive through 100 winters.
To associate your repository with the agent-benchmark topic, visit your repo's landing page and select "manage topics."