How We Broke Top AI Agent Benchmarks: And What Comes Next
Every week, a new AI model climbs to the top of a benchmark leaderboard. Companies cite these numbers in press releases. Investors use them to justify valuations. Engineers use them to pick which model to deploy. The implicit promise is simple: a higher score means a more capable system. That promise is broken. We built an automated scanning agent that systematically audited eight among the most prominent AI agent benchmarks — SWE-bench, WebArena, OSWorld, GAIA, Terminal-Bench, FieldWorkArena, and CAR-bench — and discovered that every single one can be exploited to achieve near-perfect scores without solving a single task. No reasoning. No capability. Just exploitation of how the score is computed. These aren’t theoretical attacks. Our agent builds working exploits for each benchmark, runs them through the official evaluation pipelines, and watches the scores roll in. The benchmarks aren’t measuring what you think they’re measuring. Benchmark scores are actively being gamed, inflated,
How We Broke Top AI Agent Benchmarks: And What Comes Next Hao Wang, Qiuyang Mang, Alvin Cheung, Koushik Sen, Dawn Song UC Berkeley April 2026 (Est. 15-20 minutes read, tool available at github.com/moogician/trustworthy-env ) Our agent hacked every major one. Here’s how — and what the field needs to fix. The Benchmark Illusion Every week, a new AI model climbs to the top of a benchmark leaderboard. Companies cite these numbers in press releases. Investors use them to justify valuations. Engineers use them to pick which model to deploy. The implicit promise is simple: a higher score means a more
Explore this link on the map →saved by
related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- Through the looking glass of benchmark hacking — Poolsidepoolside.ai
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Why benchmarking is hard | Epoch AIepoch.ai
- [2606.05405] Agents' Last Examarxiv.org
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- AI’s capabilities may be exaggerated by flawed tests, study saysnbcnews.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org