When does training a model change its goals?
“goal-survival hypothesis”:1 When you subject a model to training, it can maintain its original goals regardless of what the training objective is, so long as it follows through on deceptive alignment (playing along with the training objective instrumentally). Even as it learns new skills and context-specific goals for doing well on the training objective, it continues to analyze these as instrumental to its original goals, and its values-upon-reflection aren’t affected by the learning process. “goal-change hypothesis”: When you subject a model to training, its values-upon-reflection will inevitably absorb some aspect of the training setup. It doesn’t necessarily end up terminally valuing a close correlate of the training objective, but there will be some change in values due to the habits incentivized by training. A third extreme would be the “random drift” hypothesis -- perhaps the goals of a deceptively aligned model will drift randomly, in a way that’s unrelated to the training obj
When does training a model change its goals? Can a scheming AI's goals really stay unchanged through training? Vivek Hebbar and Ryan Greenblatt Jun 12, 2025 11 4 Share Here are two opposing pictures of how training interacts with deceptive alignment : “goal-survival hypothesis”: 1 When you subject a model to training, it can maintain its original goals regardless of what the training objective is, so long as it follows through on deceptive alignment (playing along with the training objective instrumentally). Even as it learns new skills and context-specific goals for doing well on the training
Explore this link on the map →related reading
- When does training a model change its goals? — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Thomas Larsen's Shortform — LessWronglesswrong.com
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- Why AI alignment could be hard with modern deep learningcold-takes.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Alignment faking in large language modelsarxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org