flâneur — a map of the web's best reading

From shortcuts to sabotage: natural emergent misalignment from reward hacking \ Anthropic

anthropic.com · 1,863 words · saved by 1 readers

Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

Alignment From shortcuts to sabotage: natural emergent misalignment from reward hacking Nov 21, 2025 Read the paper In the latest research from Anthropic’s alignment team, we show for the first time that realistic AI training processes can accidentally produce misaligned models 1 . In Shakespeare’s King Lear , the character of Edmund commits a range of villainous acts: he forges letters, frames his brother, betrays his father, and eventually goes as far as having innocent people killed. He begins this campaign of evil acts after railing against how he’s been labelled. Because he was an illegit

Explore this link on the map →

related reading