The case for ensuring that powerful AIs are controlled — LessWrong
In this post, we argue that AI labs should ensure that powerful AIs are controlled. That is, labs should make sure that the safety measures they appl…
x The case for ensuring that powerful AIs are controlled — LessWrong Best of LessWrong 2024 AI Control Redwood Research Deceptive Alignment AI Curated 290 The case for ensuring that powerful AIs are controlled by ryan_greenblatt , Buck 24th Jan 2024 AI Alignment Forum 34 min read 74 290 Ω 114 In this post, we argue that AI labs should ensure that powerful AIs are controlled . That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures. We thin
Explore this link on the map →saved by
- Elizabeth Qiu
- Shubham Chandel
- Asher P
- Lydia Nottingham
- Emil Ryd
- Jason Hausenloy
- Jo J.
- Dillon Nguyen
- Anastasia Wei
related reading
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- The case for ensuring that powerful AIs are controlled — AI Alignment Forumalignmentforum.org
- The Case Against AI Control Research — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- Introduction to AI Control - by Sarah - BlueDot Impactblog.bluedot.org
- Catching AIs red-handedblog.redwoodresearch.org
- How can we solve diffuse threats like research sabotage with AI control? — LessWronglesswrong.com
- Introduction to AI Control - by Sarah - BlueDot Impactbluedot.org