Are AI benchmarks doomed? - by Anson Ho and Greg Burnham
In this episode, Greg Burnham and Tom Adamczewski join Anson Ho to push back on benchmark pessimism and dig into what the next generation of AI benchmarks could look like.
Epoch After Hours Are AI benchmarks doomed? In this episode, Greg Burnham and Tom Adamczewski join Anson Ho to push back on benchmark pessimism and dig into what the next generation of AI benchmarks could look like. Anson Ho and Greg Burnham May 01, 2026 16 2 3 Share Greg Burnham leads Epoch’s benchmarking team. Tom Adamczewski is a senior research engineer who develops new benchmarks, including MirrorCode. Topics we cover: why benchmark saturation isn’t as alarming as it seems, how AI can speed up benchmark development, the benchmark-reality gap, whether an AGI benchmark can exist, and the ro
Explore this link on the map →saved by
related reading
- My picture of the present in AI — LessWronglesswrong.com
- Economics and Transformative AI | Tom Cunningham – Tom Cunninghamtecunningham.github.io
- After Automation | Everyevery.to
- RIP Classic Reasoning Benchmarks. What’s Next?epochai.substack.com
- AI’s capabilities may be exaggerated by flawed tests, study saysnbcnews.com
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- [2606.05405] Agents' Last Examarxiv.org
- Mapping global dynamics of benchmark creation and saturation in artificial intelligence | Nature Communicationsnature.com
- AI progress is about to speed up | Epoch AIepoch.ai
- Import AI 455: AI systems are about to start building themselves.importai.substack.com
- Import AIjack-clark.net