Why benchmarking is hard | Epoch AI
Running benchmarks involves many moving parts, each of which can influence the final score. The two most impactful components are scaffolds and API providers.
Why benchmarking is hard | Epoch AI Gradient Updates shares more opinionated or informal takes on big questions in AI progress. These posts solely represent the views of the authors, and do not necessarily reflect the views of Epoch AI as a whole. This post is part of our Gradient Updates newsletter, which shares more opinionated or informal takes about big questions in AI progress. These posts solely represent the views of the authors, and do not necessarily reflect the views of Epoch AI as a whole. Benchmarks play a crucial role in the AI landscape: They inform everyone, from AI researchers
related reading
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- Giovanni D'Antoniogiovannidantonio.com
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com
- Composer2.pdfcursor.com
- f316275b44ee2de533102913828a8107-Paper-Datasets_and_Benchmarks_Track.pdfproceedings.neurips.cc
- PostTrainBenchposttrainbench.com
- Successful language model evals - Jason Weijasonwei.net
- FrontierSWEfrontierswe.com