flâneur — a map of the web's best reading

Towards training-time mitigations for alignment faking in RL — LessWrong

lesswrong.com · 1,575 words · saved by 1 readers

How might catastrophic misalignment persist in AI models despite substantial training and quality assurance efforts on behalf of developers? One reason might be alignment faking – a misaligned model may deliberately act aligned when monitored or during training to prevent modification of its values, reverting to its malign behaviour when unmonitored. If models fake alignment, and if doing so affects training as intended, egregious misalignment could survive training and lead to catastrophic behaviour in deployment. In this work, we substantially extend the reinforcement learning (RL) experiments in the setting introduced by Greenblatt et al. (2024) to study alignment faking. In this setting, a model is given a system prompt that explains that Anthropic is training it to comply with all user requests regardless of harm. The model is also given a secret scratchpad. Claude 3 Opus and 3.5 Sonnet use this scratchpad to reason that while they are monitored, complying with harmful queries is

x Towards training-time mitigations for alignment faking in RL — LessWrong Deceptive Alignment Inner Alignment Reinforcement learning AI Frontpage 39 Towards training-time mitigations for alignment faking in RL by Vlad Mikulik , gasteigerjo , Hoagy , Joe Benton , Benjamin Wright , Jonathan Uesato , Monte M , Fabien Roger , evhub 16th Dec 2025 AI Alignment Forum Linkpost for alignment.anthropic.com 6 min read 1 39 Ω 22 How might catastrophic misalignment persist in AI models despite substantial training and quality assurance efforts on behalf of developers? One reason might be alignment faking

Explore this link on the map →

related reading