The case for ensuring that powerful AIs are controlled — AI Alignment Forum
In this post, we argue that AI labs should ensure that powerful AIs are controlled. That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures. We think no fundamental research breakthroughs are required for labs to implement safety measures that meet our standard for AI control for early transformatively useful AIs; we think that meeting our standard would substantially reduce the risks posed by intentional subversion. We recently released a paper in which we explored AI control in a simple programming setting. In that paper, we demonstrated a methodology for evaluating whether any given set of safety measures can control a model. We also introduced several techniques that let us improve the level of safety substantially above baselines without substantially sacrificing performance. In this post, we'll describe AI control in a broade
x The case for ensuring that powerful AIs are controlled — AI Alignment Forum Best of LessWrong 2024 Redwood Research AI Control Deceptive Alignment AI Curated 115 The case for ensuring that powerful AIs are controlled by ryan_greenblatt , Buck 24th Jan 2024 34 min read 74 115 In this post, we argue that AI labs should ensure that powerful AIs are controlled . That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures. We think no fundamental
Explore this link on the map →related reading
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- The Case Against AI Control Research — LessWronglesswrong.com
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Introduction to AI Control - by Sarah - BlueDot Impactblog.bluedot.org
- Introduction to AI Control - by Sarah - BlueDot Impactbluedot.org
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org