flâneur — a map of the web's best reading

AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forum

alignmentforum.org · 4,505 words · saved by 2 readers

We’ve released a paper, AI Control: Improving Safety Despite Intentional Subversion. This paper explores techniques that prevent AI catastrophes even if AI instances are colluding to subvert the safety techniques. In this post: The next post in this sequence (which we’ll release in the coming weeks) discusses what we mean by AI control and argues that it is a promising methodology for reducing risk from scheming models. Here’s the abstract of the paper: As large language models (LLMs) become more powerful and are deployed more autonomously, it will be increasingly important to prevent them from causing harmful outcomes. Researchers have investigated a variety of safety techniques for this purpose, e.g. using models to review the outputs of other models, or red-teaming techniques to surface subtle failure modes. However, researchers have not evaluated whether such techniques still ensure safety if the model is itself intentionally trying to subvert them. In this paper, we develop and ev

x AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forum Best of LessWrong 2023 AI Control Deceptive Alignment Redwood Research AI Alignment Intro Materials AI Frontpage 102 AI Control: Improving Safety Despite Intentional Subversion by Buck , Fabien Roger , ryan_greenblatt , Kshitij Sachan 13th Dec 2023 12 min read 26 102 We’ve released a paper, AI Control: Improving Safety Despite Intentional Subversion . This paper explores techniques that prevent AI catastrophes even if AI instances are colluding to subvert the safety techniques. In this post: We summarize the pap

Explore this link on the map →

saved by

related reading