Towards training-time mitigations for alignment faking in RL — LessWrong
How might catastrophic misalignment persist in AI models despite substantial training and quality assurance efforts on behalf of developers? One reason might be alignment faking – a misaligned model may deliberately act aligned when monitored or during training to prevent modification of its values, reverting to its malign behaviour when unmonitored. If models fake alignment, and if doing so affects training as intended, egregious misalignment could survive training and lead to catastrophic behaviour in deployment. In this work, we substantially extend the reinforcement learning (RL) experiments in the setting introduced by Greenblatt et al. (2024) to study alignment faking. In this setting, a model is given a system prompt that explains that Anthropic is training it to comply with all user requests regardless of harm. The model is also given a secret scratchpad. Claude 3 Opus and 3.5 Sonnet use this scratchpad to reason that while they are monitored, complying with harmful queries is
x Towards training-time mitigations for alignment faking in RL — LessWrong Deceptive Alignment Inner Alignment Reinforcement learning AI Frontpage 39 Towards training-time mitigations for alignment faking in RL by Vlad Mikulik , gasteigerjo , Hoagy , Joe Benton , Benjamin Wright , Jonathan Uesato , Monte M , Fabien Roger , evhub 16th Dec 2025 AI Alignment Forum Linkpost for alignment.anthropic.com 6 min read 1 39 Ω 22 How might catastrophic misalignment persist in AI models despite substantial training and quality assurance efforts on behalf of developers? One reason might be alignment faking
Explore this link on the map →related reading
- Alignment Faking Mitigationsalignment.anthropic.com
- Alignment faking in large language modelsarxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AIs Will Increasingly Fake Alignment - by Zvi Mowshowitzthezvi.substack.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- [2412.14093] Alignment faking in large language modelsarxiv.org
- [2412.14093] Alignment faking in large language modelsarxiv.org
- “Alignment Faking” frame is somewhat fake — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net