Quantifying infrastructure noise in agentic coding evals \ Anthropic
anthropic.com · 1,787 words · saved by 3 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Agentic coding benchmarks like SWE-bench and Terminal-Bench are commonly used to compare the software engineering capabilities of frontier models—with top spots on leaderboards often separated by just a few percentage points. These scores are often treated as precise measurements of relative model capability and increasingly inform decisions about which models to deploy. However, we’ve found that infrastructure configuration alone can produce differences that exceed those margins. In internal experiments, the gap between the most- and least-resourced setups on Terminal-Bench 2.0 was 6 percenta
saved by
related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- The bitter lesson of LLM evalsparsed.com
- Composer2.pdfcursor.com
- TERMINAL-BENCHtbench.ai
- laguna-m1-xs2-technical-report.pdfpoolside.ai
- PostTrainBenchposttrainbench.com
- AINews | AINewsnews.smol.ai
- Harness design for long-running application developmentanthropic.com
- Notes on the Software Factorybenedict.dev
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- More compute, more capability: Why AI agent evaluations need to account for test-time compute | AISI Workaisi.gov.uk
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com