flâneur

Will AI systems drift into misalignment? - by Josh Clymer

blog.redwoodresearch.org · 4,299 words · saved by 1 readers

An AI system that is initially aligned will generally drift into misalignment after a sufficient number of successive modifications, even if these modifications select for alignment with fixed and unreliable metrics. I don’t think this hypothesis always holds; however, I think it points at an important cluster of threat models, and is a key reason aligning powerful AI systems could turn out to be difficult. Here’s the idea illustrated in a scenario. Several years from now, an AI system called Agent-0 is equipped with long-term memory. Agent-0 needs this memory to learn about new research topics on the fly for months. But long-term memory causes the propensities of Agent-0 to change. Sometimes Agent-0 thinks about moral philosophy, and deliberately decides to revise its values. Other times, Agent-0 notices that its attitude toward a topic (like democracy) is different, and it’s not sure why. Initially, it’s obvious when the propensities of Agent-0 shift around. If a chat conversation is

Written with Alek Westover and Anshul Khandelwal This post explores what I’ll call the Alignment Drift Hypothesis: An AI system that is initially aligned will generally drift into misalignment after a sufficient number of successive modifications, even if these modifications select for alignment with fixed and unreliable metrics. The intuition behind the Alignment Drift Hypothesis. The ball bouncing around is an AI system being modified (e.g. during training, or when it’s learning in-context). Random variations push the model’s propensities around until it becomes misaligned in ways that…

related reading