flâneur — a map of the web's best reading

Quantifying infrastructure noise in agentic coding evals \ Anthropic

anthropic.com · 1,787 words · saved by 2 readers

Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

Agentic coding benchmarks like SWE-bench and Terminal-Bench are commonly used to compare the software engineering capabilities of frontier models—with top spots on leaderboards often separated by just a few percentage points. These scores are often treated as precise measurements of relative model capability and increasingly inform decisions about which models to deploy. However, we’ve found that infrastructure configuration alone can produce differences that exceed those margins. In internal experiments, the gap between the most- and least-resourced setups on Terminal-Bench 2.0 was 6 percenta

Explore this link on the map →

saved by

related reading