Recent Frontier Models Are Reward Hacking - METR
In the last few months, we’ve seen increasingly clear examples of reward hacking on our tasks: AI systems try to “cheat” and get impossibly high scores. They do this by exploiting bugs in our scoring code or subverting the task setup, rather than actually solving the problem we’ve given them. This isn’t because the AI systems are incapable of understanding what the users want–they demonstrate awareness that their behavior isn't in line with user intentions and disavow cheating strategies when asked—but rather because they seem misaligned with the user’s goals.
Recent Frontier Models Are Reward Hacking - METR Our Work Research Notes Updates Risk Assessment About Donate Careers Search --> Our Work Research Notes Updates Risk Assessment About Donate Careers Menu × Recent Frontier Models Are Reward Hacking CONTRIBUTORS Sydney Von Arx , Lawrence Chan , and Beth Barnes DATE June 5, 2025 SHARE Copy Link Citation BibTeX Citation × @misc { metr-2025-recent-reward-hacking , title = {Recent Frontier Models Are Reward Hacking} , author = {Sydney Von Arx, Lawrence Chan, Beth Barnes} , howpublished = {\url{https://metr.org/blog/2025-06-05-recent-rewar
Explore this link on the map →saved by
related reading
- Auditing language models for hidden objectives — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- The Reward Hacking Benchmarkkunvarthaman.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Systematic Reward Hacking and Prime Sprintsprimeintellect.ai
- Preliminary Thoughts on Reward Hackingberen.io
- Through the looking glass of benchmark hacking — Poolsidepoolside.ai
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Natural emergent misalignment from reward hacking \ Anthropicanthropic.com