When does training a model change its goals? — LessWrong
A third extreme would be the “random drift” hypothesis -- perhaps the goals of a deceptively aligned model will drift randomly, in a way that’s unrelated to the training objective. A closely related question is “when do instrumental goals become terminal?” The goal-survival hypothesis would imply that instrumental goals generally don’t become terminal, while the goal-change hypothesis is most compatible with a world where instrumental goals often become terminal. The question of goal-survival vs. goal-change comes up in many places we care about: Which hypothesis is true in a given setting will surely depend on the capabilities of the model and on the nature and duration of training. However, it would be nice to know which way reality leans overall. The results in sleeper agents, alignment faking, and other works are highly relevant to this question -- the next section of this post reviews the evidence from those papers. The final section of this post lays out some slightly unhinge
x When does training a model change its goals? — LessWrong Deceptive Alignment AI Frontpage 79 When does training a model change its goals? by Vivek Hebbar , ryan_greenblatt 12th Jun 2025 AI Alignment Forum 18 min read 3 79 Ω 42 Here are two opposing pictures of how training interacts with deceptive alignment : “goal-survival hypothesis”: [1] When you subject a model to training, it can maintain its original goals regardless of what the training objective is, so long as it follows through on deceptive alignment (playing along with the training objective instrumentally). Even as it learns new s
Explore this link on the map →related reading
- When does training a model change its goals?redwoodresearch.substack.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Thomas Larsen's Shortform — LessWronglesswrong.com
- Why AI alignment could be hard with modern deep learningcold-takes.com
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- Teaching Claude why \ Anthropicanthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Models Don't "Get Reward" — LessWronglesswrong.com