Are AI benchmarks doomed? - by Anson Ho and Greg Burnham
In this episode, Greg Burnham and Tom Adamczewski join Anson Ho to push back on benchmark pessimism and dig into what the next generation of AI benchmarks could look like.
Epoch After Hours Are AI benchmarks doomed? In this episode, Greg Burnham and Tom Adamczewski join Anson Ho to push back on benchmark pessimism and dig into what the next generation of AI benchmarks could look like. Anson Ho and Greg Burnham May 01, 2026 16 2 3 Share Greg Burnham leads Epoch’s benchmarking team. Tom Adamczewski is a senior research engineer who develops new benchmarks, including MirrorCode. Topics we cover: why benchmark saturation isn’t as alarming as it seems, how AI can speed up benchmark development, the benchmark-reality gap, whether an AGI benchmark can exist, and the ro
saved by
related reading
- My picture of the present in AI — LessWronglesswrong.com
- Giovanni D'Antoniogiovannidantonio.com
- After Automation | Everyevery.to
- RIP Classic Reasoning Benchmarks. What’s Next?epochai.substack.com
- AI’s capabilities may be exaggerated by flawed tests, study saysnbcnews.com
- On measuring AInikilravi.substack.com
- f316275b44ee2de533102913828a8107-Paper-Datasets_and_Benchmarks_Track.pdfproceedings.neurips.cc
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- [2606.05405] Agents' Last Examarxiv.org
- Measuring AI Ability to Complete Long Software Tasksarxiv.org
- AI progress is about to speed up | Epoch AIepoch.ai
- Measuring AI Ability to Complete Long Tasks - METRmetr.org