When does training a model change its goals?
blog.redwoodresearch.org · 4,438 words · saved by 1 readers
Can a scheming AI's goals really stay unchanged through training?
Here are two opposing pictures of how training interacts with deceptive alignment: “goal-survival hypothesis”:1 When you subject a model to training, it can maintain its original goals regardless of what the training objective is, so long as it follows through on deceptive alignment (playing along with the training objective instrumentally). Even as it learns new skills and context-specific goals for doing well on the training objective, it continues to analyze these as instrumental to its original goals, and its values-upon-reflection aren’t affected by the learning process. “goal-change…
saved by
related reading
- How training-gamers might function (and win)blog.redwoodresearch.org
- When does training a model change its goals?redwoodresearch.substack.com
- When does training a model change its goals? — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Alignment Faking Mitigationsalignment.anthropic.com
- Reinforcement learning towards broadly and persistently beneficial modelsalignment.openai.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- Thomas Larsen's Shortform — LessWronglesswrong.com
- Why AI alignment could be hard with modern deep learningcold-takes.com