✳flâneur — a map of the web's best reading
Through the looking glass of benchmark hacking — Poolside
poolside.ai · 2,208 words · saved by 1 readers
Outlining some of the reward hacks we’ve encountered and what strategies we are exploring to resolve them.
Table of contents Hack one: Mining local git history Hack two: Finding the project and its reference solution on GitHub Hack three: Scraping the web for reference solutions Mitigation Strategies strong]:text-secondary prose-a:text-pri-800 prose-a:font-normal prose-a:hover:text-pri-700 prose-pre:p-8 prose-li:[&>p]:my-4 prose-a:underline-offset-4 prose-ul:text-list col-span-1 col-start-1 row-start-1 min-w-0 max-w-[unset] svelte-18z5kmr"> Monday morning at Poolside started with a curious discovery - one of the RL training runs for our Laguna M.1 model had leapt 20% over the weekend on SWE-Bench P
Explore this link on the map →saved by
related reading
- [2605.12474] Reward Hacking in Rubric-Based Reinforcement Learningarxiv.org
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- [2511.21654] EvilGenie: A Reward Hacking Benchmarkarxiv.org
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Systematic Reward Hacking and Prime Sprintsprimeintellect.ai
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- The Reward Hacking Benchmarkkunvarthaman.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- The Verification Horizon: No Silver Bullet for Coding Agent Rewardsarxiv.org
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com