flâneur

Every Benchmark is Broken - Jonathan’s Substack

jonathanpgabor.substack.com · 1,639 words · saved by 1 readers

feat. RE-Bench, HLE, LCB-Pro, Terminal-Bench 2, and more!

Last June, METR caught o3 reward hacking on its RE-Bench and HCAST benchmarks. In a particularly humorous case, o3, when tasked with optimizing a kernel, decided to “shrink the notion of time as seen by the scorer”. The development of Humanity’s Last Exam involved “over 1,000 subject-matter experts” and $500,000 in prizes. However, after its release, researchers at FutureHouse discovered “about 30% of chemistry/biology answers are likely wrong”. LiveCodeBench Pro is a competitive programming benchmark developed by “a group of medalists in international algorithmic contests”. Their paper…

saved by

related reading