✳flâneur — a map of the web's best reading
Quantifying infrastructure noise in agentic coding evals \ Anthropic
anthropic.com · 1,787 words · saved by 2 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Agentic coding benchmarks like SWE-bench and Terminal-Bench are commonly used to compare the software engineering capabilities of frontier models—with top spots on leaderboards often separated by just a few percentage points. These scores are often treated as precise measurements of relative model capability and increasingly inform decisions about which models to deploy. However, we’ve found that infrastructure configuration alone can produce differences that exceed those margins. In internal experiments, the gap between the most- and least-resourced setups on Terminal-Bench 2.0 was 6 percenta
Explore this link on the map →saved by
related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Composer2.pdfcursor.com
- AI Benchmark Leaderboards & Model Evals | BenchmarkListbenchmarklist.com
- PostTrainBenchposttrainbench.com
- The bitter lesson of LLM evalsparsed.com
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Why benchmarking is hard | Epoch AIepoch.ai
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Effective harnesses for long-running agents \ Anthropicanthropic.com