Maintaining Alignment during RSI as a Feedback Control Problem
Recent advances have begun to move AI beyond pretrained amortized models and supervised learning. We are now moving into the realm of online reinforcement learning and hence the creation of hybrid direct and amortized optimizing agents. While we generally have found that purely amortized pretrained models are an easy case...
Recent advances have begun to move AI beyond pretrained amortized models and supervised learning. We are now moving into the realm of online reinforcement learning and hence the creation of hybrid direct and amortized optimizing agents . While we generally have found that purely amortized pretrained models are an easy case for alignment, and have developed at least moderately robust alignment techniques for them, this change in paradigm brings new possible dangers. Looking even further ahead, as we move towards agents that are capable of continual online learning and ultimately recursive self
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- I think alignment work is more promising than control work — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- The Case Against AI Control Research — LessWronglesswrong.com
- Reading Listblog.redwoodresearch.org
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Introduction to AI Control - by Sarah - BlueDot Impactblog.bluedot.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- How useful is AI control? @ Trackstracks.xlabtracks.workers.dev