When does training a model change its goals?
“goal-survival hypothesis”:1 When you subject a model to training, it can maintain its original goals regardless of what the training objective is, so long as it follows through on deceptive alignment (playing along with the training objective instrumentally). Even as it learns new skills and context-specific goals for doing well on the training objective, it continues to analyze these as instrumental to its original goals, and its values-upon-reflection aren’t affected by the learning process. “goal-change hypothesis”: When you subject a model to training, its values-upon-reflection will inevitably absorb some aspect of the training setup. It doesn’t necessarily end up terminally valuing a close correlate of the training objective, but there will be some change in values due to the habits incentivized by training. A third extreme would be the “random drift” hypothesis -- perhaps the goals of a deceptively aligned model will drift randomly, in a way that’s unrelated to the training obj
When does training a model change its goals? Can a scheming AI's goals really stay unchanged through training? Vivek Hebbar and Ryan Greenblatt Jun 12, 2025 11 4 Share Here are two opposing pictures of how training interacts with deceptive alignment : “goal-survival hypothesis”: 1 When you subject a model to training, it can maintain its original goals regardless of what the training objective is, so long as it follows through on deceptive alignment (playing along with the training objective instrumentally). Even as it learns new skills and context-specific goals for doing well on the training
related reading
- When does training a model change its goals?blog.redwoodresearch.org
- When does training a model change its goals? — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Reinforcement learning towards broadly and persistently beneficial modelsalignment.openai.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Teaching Claude Whyalignment.anthropic.com
- Thomas Larsen's Shortform — LessWronglesswrong.com
- Deep Deceptiveness — LessWronglesswrong.com
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org