flâneur — a map of the web's best reading

Buck Shlegeris on controlling AI that wants to take over – so we can use it anyway - 80,000 Hours

80000hours.org · 38,796 words · saved by 1 readers

Another thing you might do once you’ve caught your AI trying to escape is let the AI think it succeeded, and see what it does. Presumably when the AI got caught, it was in the midst of taking some actions that it thought would let it launch a rogue deployment. If your model knows a bunch of security vulnerabilities in your software, it might start deploying all those things. — Buck Shlegeris Most AI safety conversations centre on alignment: ensuring AI systems share our values and goals. But despite progress, we’re unlikely to know we’ve solved the problem before the arrival of human-level and superhuman systems in as little as three years. So some are developing a backup plan to safely deploy models we fear are actively scheming to harm us — so-called “AI control.” While this may sound mad, given the reluctance of AI companies to delay deploying anything they train, not developing such techniques is probably even crazier. Today’s guest — Buck Shlegeris, CEO of Redwood Research — has s

Buck Shlegeris on controlling AI that wants to take over – so we can use it anyway | 80,000 Hours Search for: On this page: Introduction 1 Highlights 2 Articles, books, and other media discussed in the show 3 Transcript 3.1 Cold open [00:00:00] 3.2 Who's Buck Shlegeris? [00:01:25] 3.3 What's AI control? [00:01:51] 3.4 Why is AI control hot now? [00:05:46] 3.5 Detecting human vs AI spies [00:10:44] 3.6 Acute vs chronic AI betrayal [00:15:41] 3.7 How to catch AIs trying to escape [00:18:10] 3.8 The cheapest AI control techniques [00:33:18] 3.9 Can we get untrusted models to do trusted work? [00:

Explore this link on the map →

related reading