AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forum
We’ve released a paper, AI Control: Improving Safety Despite Intentional Subversion. This paper explores techniques that prevent AI catastrophes even if AI instances are colluding to subvert the safety techniques. In this post: The next post in this sequence (which we’ll release in the coming weeks) discusses what we mean by AI control and argues that it is a promising methodology for reducing risk from scheming models. Here’s the abstract of the paper: As large language models (LLMs) become more powerful and are deployed more autonomously, it will be increasingly important to prevent them from causing harmful outcomes. Researchers have investigated a variety of safety techniques for this purpose, e.g. using models to review the outputs of other models, or red-teaming techniques to surface subtle failure modes. However, researchers have not evaluated whether such techniques still ensure safety if the model is itself intentionally trying to subvert them. In this paper, we develop and ev
x AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forum Best of LessWrong 2023 AI Control Deceptive Alignment Redwood Research AI Alignment Intro Materials AI Frontpage 102 AI Control: Improving Safety Despite Intentional Subversion by Buck , Fabien Roger , ryan_greenblatt , Kshitij Sachan 13th Dec 2023 12 min read 26 102 We’ve released a paper, AI Control: Improving Safety Despite Intentional Subversion . This paper explores techniques that prevent AI catastrophes even if AI instances are colluding to subvert the safety techniques. In this post: We summarize the pap
Explore this link on the map →saved by
related reading
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- The Case Against AI Control Research — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- 2312.06942arxiv.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionar5iv.labs.arxiv.org
- AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org