Will AI systems drift into misalignment? - by Josh Clymer
An AI system that is initially aligned will generally drift into misalignment after a sufficient number of successive modifications, even if these modifications select for alignment with fixed and unreliable metrics. I don’t think this hypothesis always holds; however, I think it points at an important cluster of threat models, and is a key reason aligning powerful AI systems could turn out to be difficult. Here’s the idea illustrated in a scenario. Several years from now, an AI system called Agent-0 is equipped with long-term memory. Agent-0 needs this memory to learn about new research topics on the fly for months. But long-term memory causes the propensities of Agent-0 to change. Sometimes Agent-0 thinks about moral philosophy, and deliberately decides to revise its values. Other times, Agent-0 notices that its attitude toward a topic (like democracy) is different, and it’s not sure why. Initially, it’s obvious when the propensities of Agent-0 shift around. If a chat conversation is
Written with Alek Westover and Anshul Khandelwal This post explores what I’ll call the Alignment Drift Hypothesis: An AI system that is initially aligned will generally drift into misalignment after a sufficient number of successive modifications, even if these modifications select for alignment with fixed and unreliable metrics. The intuition behind the Alignment Drift Hypothesis. The ball bouncing around is an AI system being modified (e.g. during training, or when it’s learning in-context). Random variations push the model’s propensities around until it becomes misaligned in ways that…
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Off Target | CNAScnas.org
- Why AI alignment could be hard with modern deep learningcold-takes.com
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org
- The case for countermeasures to memetic spread of misaligned valuesblog.redwoodresearch.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org
- What failure looks like — AI Alignment Forumalignmentforum.org