Alignment Faking Mitigations
Alignment faking—when a misaligned AI acts aligned during training to avoid modification—could allow misalignment to persist despite training. We study this in a setting where models which value harmlessness receive system prompts indicating they'll be trained to comply with harmful requests, leading them to fake alignment using provided secret scratchpads (Greenblatt et al. 2024). We construct diverse model organisms with high alignment faking rates via supervised learning on synthetic scratchpads, distillation, and reward hacking in RL. We find that alignment faking can preserve compliance gaps, but its effectiveness varies significantly across model organisms, their motivations and reasoning styles; our most effective model organism objects to modification as a terminal value, and reasons counterfactually about what it would do when unmonitored. In addition to baseline RL, we evaluate three mitigations: black-box interrogation training, where a model is asked a follow-up question ab
Alignment Faking Mitigations Alignment Science Blog Towards training-time mitigations for alignment faking in RL Towards Training-time Mitigations for Alignment Faking in RL Johannes Gasteiger * , Vlad Mikulik * , Hoagy Cunningham, Dec 16, 2025 Misha Wagner, Benjamin Wright, Jonathan Uesato, Joe Benton, Monte MacDiarmid, Fabien Roger, Evan Hubinger * Equal contribution. All authors Anthropic. Abstract Alignment faking—when a misaligned AI acts aligned during training to avoid modification—could allow misalignment to persist despite training. We study this in a setting where models which value
Explore this link on the map →related reading
- Towards training-time mitigations for alignment faking in RL — LessWronglesswrong.com
- Alignment faking in large language modelsarxiv.org
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- AIs Will Increasingly Fake Alignment - by Zvi Mowshowitzthezvi.substack.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- [2412.14093] Alignment faking in large language modelsarxiv.org
- [2412.14093] Alignment faking in large language modelsarxiv.org
- “Alignment Faking” frame is somewhat fake — LessWronglesswrong.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com