From shortcuts to sabotage: natural emergent misalignment from reward hacking \ Anthropic
anthropic.com · 1,863 words · saved by 1 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Alignment From shortcuts to sabotage: natural emergent misalignment from reward hacking Nov 21, 2025 Read the paper In the latest research from Anthropic’s alignment team, we show for the first time that realistic AI training processes can accidentally produce misaligned models 1 . In Shakespeare’s King Lear , the character of Edmund commits a range of villainous acts: he forges letters, frames his brother, betrays his father, and eventually goes as far as having innocent people killed. He begins this campaign of evil acts after railing against how he’s been labelled. Because he was an illegit
related reading
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Training a Misaligned Reward Seekeralignment.anthropic.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Teaching Claude Whyalignment.anthropic.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Alignment Faking Mitigationsalignment.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com