flâneur — a map of the web's best reading

When does training a model change its goals?

redwoodresearch.substack.com · 4,551 words · saved by 1 readers

“goal-survival hypothesis”:1 When you subject a model to training, it can maintain its original goals regardless of what the training objective is, so long as it follows through on deceptive alignment (playing along with the training objective instrumentally). Even as it learns new skills and context-specific goals for doing well on the training objective, it continues to analyze these as instrumental to its original goals, and its values-upon-reflection aren’t affected by the learning process. “goal-change hypothesis”: When you subject a model to training, its values-upon-reflection will inevitably absorb some aspect of the training setup. It doesn’t necessarily end up terminally valuing a close correlate of the training objective, but there will be some change in values due to the habits incentivized by training. A third extreme would be the “random drift” hypothesis -- perhaps the goals of a deceptively aligned model will drift randomly, in a way that’s unrelated to the training obj

When does training a model change its goals? Can a scheming AI's goals really stay unchanged through training? Vivek Hebbar and Ryan Greenblatt Jun 12, 2025 11 4 Share Here are two opposing pictures of how training interacts with deceptive alignment : “goal-survival hypothesis”: 1 When you subject a model to training, it can maintain its original goals regardless of what the training objective is, so long as it follows through on deceptive alignment (playing along with the training objective instrumentally). Even as it learns new skills and context-specific goals for doing well on the training

Explore this link on the map →

related reading