flâneur — a map of the web's best reading

AI Control: Improving Safety Despite Intentional Subversion — LessWrong

lesswrong.com · 7,286 words · saved by 4 readers

We’ve released a paper, AI Control: Improving Safety Despite Intentional Subversion. This paper explores techniques that prevent AI catastrophes even…

x AI Control: Improving Safety Despite Intentional Subversion — LessWrong Best of LessWrong 2023 AI Control Deceptive Alignment Redwood Research AI Alignment Intro Materials AI Frontpage 240 AI Control: Improving Safety Despite Intentional Subversion by Buck , Fabien Roger , ryan_greenblatt , Kshitij Sachan 13th Dec 2023 AI Alignment Forum 12 min read 26 240 Ω 102 We’ve released a paper, AI Control: Improving Safety Despite Intentional Subversion . This paper explores techniques that prevent AI catastrophes even if AI instances are colluding to subvert the safety techniques. In this post: We s

Explore this link on the map →

saved by

related reading