flâneur

When does training a model change its goals?

blog.redwoodresearch.org · 4,438 words · saved by 1 readers

Can a scheming AI's goals really stay unchanged through training?

Here are two opposing pictures of how training interacts with deceptive alignment: “goal-survival hypothesis”:1 When you subject a model to training, it can maintain its original goals regardless of what the training objective is, so long as it follows through on deceptive alignment (playing along with the training objective instrumentally). Even as it learns new skills and context-specific goals for doing well on the training objective, it continues to analyze these as instrumental to its original goals, and its values-upon-reflection aren’t affected by the learning process. “goal-change…

saved by

related reading