Buck Shlegeris on controlling AI that wants to take over – so we can use it anyway - 80,000 Hours
Another thing you might do once you’ve caught your AI trying to escape is let the AI think it succeeded, and see what it does. Presumably when the AI got caught, it was in the midst of taking some actions that it thought would let it launch a rogue deployment. If your model knows a bunch of security vulnerabilities in your software, it might start deploying all those things. — Buck Shlegeris Most AI safety conversations centre on alignment: ensuring AI systems share our values and goals. But despite progress, we’re unlikely to know we’ve solved the problem before the arrival of human-level and superhuman systems in as little as three years. So some are developing a backup plan to safely deploy models we fear are actively scheming to harm us — so-called “AI control.” While this may sound mad, given the reluctance of AI companies to delay deploying anything they train, not developing such techniques is probably even crazier. Today’s guest — Buck Shlegeris, CEO of Redwood Research — has s
Buck Shlegeris on controlling AI that wants to take over – so we can use it anyway | 80,000 Hours Search for: On this page: Introduction 1 Highlights 2 Articles, books, and other media discussed in the show 3 Transcript 3.1 Cold open [00:00:00] 3.2 Who's Buck Shlegeris? [00:01:25] 3.3 What's AI control? [00:01:51] 3.4 Why is AI control hot now? [00:05:46] 3.5 Detecting human vs AI spies [00:10:44] 3.6 Acute vs chronic AI betrayal [00:15:41] 3.7 How to catch AIs trying to escape [00:18:10] 3.8 The cheapest AI control techniques [00:33:18] 3.9 Can we get untrusted models to do trusted work? [00:
Explore this link on the map →related reading
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Catching AIs red-handedblog.redwoodresearch.org
- AI 2027ai-2027.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Catching AIs red-handed — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- The Case Against AI Control Research — LessWronglesswrong.com
- gdm-ai-control-roadmap.pdfstorage.googleapis.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- The case for ensuring that powerful AIs are controlled — AI Alignment Forumalignmentforum.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org