✳flâneur — a map of the web's best reading
AI catastrophes and rogue deployments - by Buck Shlegeris
blog.redwoodresearch.org · 2,345 words · saved by 2 readers
It’s interesting to classify possible AI catastrophes based on whether or not they involve a "rogue deployment".
AI catastrophes and rogue deployments It’s interesting to classify possible AI catastrophes based on whether or not they involve a "rogue deployment". Buck Shlegeris Jun 03, 2024 5 Share [Thanks to Aryan Bhatt, Ansh Radhakrishnan, Adam Kaufman, Vivek Hebbar, Hanna Gabor, Justis Mills, Aaron Scher, Max Nadeau, Ryan Greenblatt, Peter Barnett, Fabien Roger, and various people at a presentation of these arguments for comments. These ideas aren’t very original to me; many of the examples of threat models are from other people.] In this post, I want to introduce the concept of a “rogue deployment” a
Explore this link on the map →saved by
related reading
- AI catastrophes and rogue deployments — AI Alignment Forumalignmentforum.org
- Catching AIs red-handedblog.redwoodresearch.org
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Three Sketches of ASL-4 Safety Case Componentsalignment.anthropic.com
- The Rogue Replication Threat Model - METRmetr.org
- gdm-ai-control-roadmap.pdfstorage.googleapis.com
- A basic systems architecture for AI agents that do autonomous research — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org