✳flâneur — a map of the web's best reading
From shortcuts to sabotage: natural emergent misalignment from reward hacking \ Anthropic
anthropic.com · 1,863 words · saved by 1 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Alignment From shortcuts to sabotage: natural emergent misalignment from reward hacking Nov 21, 2025 Read the paper In the latest research from Anthropic’s alignment team, we show for the first time that realistic AI training processes can accidentally produce misaligned models 1 . In Shakespeare’s King Lear , the character of Edmund commits a range of villainous acts: he forges letters, frames his brother, betrays his father, and eventually goes as far as having innocent people killed. He begins this campaign of evil acts after railing against how he’s been labelled. Because he was an illegit
Explore this link on the map →related reading
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com