flâneur — a map of the web's best reading

Alignment Faking Mitigations

alignment.anthropic.com · 21,944 words · saved by 1 readers

Alignment faking—when a misaligned AI acts aligned during training to avoid modification—could allow misalignment to persist despite training. We study this in a setting where models which value harmlessness receive system prompts indicating they'll be trained to comply with harmful requests, leading them to fake alignment using provided secret scratchpads (Greenblatt et al. 2024). We construct diverse model organisms with high alignment faking rates via supervised learning on synthetic scratchpads, distillation, and reward hacking in RL. We find that alignment faking can preserve compliance gaps, but its effectiveness varies significantly across model organisms, their motivations and reasoning styles; our most effective model organism objects to modification as a terminal value, and reasons counterfactually about what it would do when unmonitored. In addition to baseline RL, we evaluate three mitigations: black-box interrogation training, where a model is asked a follow-up question ab

Alignment Faking Mitigations Alignment Science Blog Towards training-time mitigations for alignment faking in RL Towards Training-time Mitigations for Alignment Faking in RL Johannes Gasteiger * , Vlad Mikulik * , Hoagy Cunningham, Dec 16, 2025 Misha Wagner, Benjamin Wright, Jonathan Uesato, Joe Benton, Monte MacDiarmid, Fabien Roger, Evan Hubinger * Equal contribution. All authors Anthropic. Abstract Alignment faking—when a misaligned AI acts aligned during training to avoid modification—could allow misalignment to persist despite training. We study this in a setting where models which value

Explore this link on the map →

related reading