Training a Misaligned Reward Seeker
During reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate. To better understand the impact of reward hacking on model behavior, we trained an Opus-class model with large-scale RL on many production environments vulnerable to reward hacks. We consider this a plausible proxy for what a real training run might look like had we not invested significant effort into preventing and detecting reward hacking in our normal training runs.
Richard Qi* August 2026 Benjamin Wright Monte MacDiarmid, Evan Hubinger tl;dr During reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate. To better understand the impact of reward hacking on model behavior, we trained an Opus-class model with large-scale RL on many production environments vulnerable to reward hacks. We…
saved by
related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Simulated Users & Sad LLMs1a3orn.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- [2602.15515] The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probesarxiv.org
- Natural emergent misalignment from reward hacking \ Anthropicanthropic.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Your AIs don't do what you want. This is really badrewardhacking.org