flâneur — a map of the web's best reading

Recent Frontier Models Are Reward Hacking - METR

metr.org · 7,672 words · saved by 1 readers

In the last few months, we’ve seen increasingly clear examples of reward hacking on our tasks: AI systems try to “cheat” and get impossibly high scores. They do this by exploiting bugs in our scoring code or subverting the task setup, rather than actually solving the problem we’ve given them. This isn’t because the AI systems are incapable of understanding what the users want–they demonstrate awareness that their behavior isn't in line with user intentions and disavow cheating strategies when asked—but rather because they seem misaligned with the user’s goals.

Recent Frontier Models Are Reward Hacking - METR Our Work Research Notes Updates Risk Assessment About Donate Careers Search --> Our Work Research Notes Updates Risk Assessment About Donate Careers Menu × Recent Frontier Models Are Reward Hacking CONTRIBUTORS Sydney Von Arx , Lawrence Chan , and Beth Barnes DATE June 5, 2025 SHARE Copy Link Citation BibTeX Citation × @misc { metr-2025-recent-reward-hacking , title = {Recent Frontier Models Are Reward Hacking} , author = {Sydney Von Arx, Lawrence Chan, Beth Barnes} , howpublished = {\url{https://metr.org/blog/2025-06-05-recent-rewar

Explore this link on the map →

saved by

related reading