Every Benchmark is Broken - Jonathan’s Substack
jonathanpgabor.substack.com · 1,639 words · saved by 1 readers
feat. RE-Bench, HLE, LCB-Pro, Terminal-Bench 2, and more!
Last June, METR caught o3 reward hacking on its RE-Bench and HCAST benchmarks. In a particularly humorous case, o3, when tasked with optimizing a kernel, decided to “shrink the notion of time as seen by the scorer”. The development of Humanity’s Last Exam involved “over 1,000 subject-matter experts” and $500,000 in prizes. However, after its release, researchers at FutureHouse discovered “about 30% of chemistry/biology answers are likely wrong”. LiveCodeBench Pro is a competitive programming benchmark developed by “a group of medalists in international algorithmic contests”. Their paper…
saved by
related reading
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering | alphaXivalphaxiv.org
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- Through the looking glass of benchmark hacking — Poolsidepoolside.ai
- [2501.01257] CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratingsarxiv.org
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Composer2.pdfcursor.com
- KernelBench v0.1 | Scaling Intelligence Lab at Stanford Universityscalingintelligence.stanford.edu
- Introducing FrontierCode | Cognitioncognition.ai
- Introducing FrontierCode | Cognitioncognition.com
- [2511.21654] EvilGenie: A Reward Hacking Benchmarkarxiv.org
- Giovanni D'Antoniogiovannidantonio.com