Why benchmarking is hard | Epoch AI
Running benchmarks involves many moving parts, each of which can influence the final score. The two most impactful components are scaffolds and API providers.
Why benchmarking is hard | Epoch AI Gradient Updates shares more opinionated or informal takes on big questions in AI progress. These posts solely represent the views of the authors, and do not necessarily reflect the views of Epoch AI as a whole. This post is part of our Gradient Updates newsletter, which shares more opinionated or informal takes about big questions in AI progress. These posts solely represent the views of the authors, and do not necessarily reflect the views of Epoch AI as a whole. Benchmarks play a crucial role in the AI landscape: They inform everyone, from AI researchers
Explore this link on the map →related reading
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com
- Composer2.pdfcursor.com
- gpt-4.pdfcdn.openai.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- PostTrainBenchposttrainbench.com
- The bitter lesson of LLM evalsparsed.com
- Challenges in evaluating AI systems \ Anthropicanthropic.com