AI catastrophes and rogue deployments — AI Alignment Forum
[Thanks to Aryan Bhatt, Ansh Radhakrishnan, Adam Kaufman, Vivek Hebbar, Hanna Gabor, Justis Mills, Aaron Scher, Max Nadeau, Ryan Greenblatt, Peter Barnett, Fabien Roger, and various people at a presentation of these arguments for comments. These ideas aren’t very original to me; many of the examples of threat models are from other people.] In this post, I want to introduce the concept of a “rogue deployment” and argue that it’s interesting to classify possible AI catastrophes based on whether or not they involve a rogue deployment. I’ll also talk about how this division interacts with the structure of a safety case, discuss two important subcategories of rogue deployment, and make a few points about how the different categories I describe here might be caused by different attackers (e.g. the AI itself, rogue lab insiders, external hackers, or multiple of these at once). Suppose you’ve developed some powerful model. (It might actually be many models; I’m just going to talk about a singu
x AI catastrophes and rogue deployments — AI Alignment Forum AI Curated 48 AI catastrophes and rogue deployments by Buck 3rd Jun 2024 10 min read 17 48 [Thanks to Aryan Bhatt, Ansh Radhakrishnan, Adam Kaufman, Vivek Hebbar, Hanna Gabor, Justis Mills, Aaron Scher, Max Nadeau, Ryan Greenblatt, Peter Barnett, Fabien Roger, and various people at a presentation of these arguments for comments. These ideas aren’t very original to me; many of the examples of threat models are from other people.] In this post, I want to introduce the concept of a “rogue deployment” and argue that it’s interesting to c
related reading
- AI catastrophes and rogue deployments - by Buck Shlegerisblog.redwoodresearch.org
- Reading Listblog.redwoodresearch.org
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Three Sketches of ASL-4 Safety Case Componentsalignment.anthropic.com
- gdm-ai-control-roadmap.pdfstorage.googleapis.com
- What failure looks like — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Frontier Risk Report (February to March 2026) - METRmetr.org
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com