flâneur

More compute, more capability: Why AI agent evaluations need to account for test-time compute | AISI Work

aisi.gov.uk · 2,083 words · saved by 1 readers

Standard evaluations cap how much compute AI agents can use. We show that raising those caps changes measured capability, the difficulty of tasks agents can solve, and how fast the frontier appears to move.

As AI agents are given more autonomy and harder tasks, underestimating their capabilities becomes more consequential. Most agent evaluations still reduce capability to a single number: a benchmark score, a pass/fail, or the length of task an agent can finish. That number hides a key design choice: how much compute the agent is allowed to spend before stopping. In March 2026, AISI published results suggesting that modest compute limits understate model capability on cyber tasks, and that the same limits may now miss substantial gains. Our Science of Evaluation team has since run frontier…

saved by

related reading