More compute, more capability: Why AI agent evaluations need to account for test-time compute | AISI Work
Standard evaluations cap how much compute AI agents can use. We show that raising those caps changes measured capability, the difficulty of tasks agents can solve, and how fast the frontier appears to move.
As AI agents are given more autonomy and harder tasks, underestimating their capabilities becomes more consequential. Most agent evaluations still reduce capability to a single number: a benchmark score, a pass/fail, or the length of task an agent can finish. That number hides a key design choice: how much compute the agent is allowed to spend before stopping. In March 2026, AISI published results suggesting that modest compute limits understate model capability on cyber tasks, and that the same limits may now miss substantial gains. Our Science of Evaluation team has since run frontier…
saved by
related reading
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com
- Metrics of Agent Ability - METRmetr.org
- My picture of the present in AI — LessWronglesswrong.com
- AI progress is about to speed up | Epoch AIepoch.ai
- A Few Things I Learned About Evals - Ryan Bloomryanbloom.xyz
- Measuring AI Ability to Complete Long Software Tasksarxiv.org
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- 2503.14499arxiv.org