Notes on handling non-concentrated failures with AI control: high level methods and different regimes — LessWrong
In this post, I'll try to explain my current understanding of the high level methods for handling non-concentrated failures with control. I'll discus…
x Notes on handling non-concentrated failures with AI control: high level methods and different regimes — LessWrong AI Control AI Frontpage 24 Notes on handling non-concentrated failures with AI control: high level methods and different regimes by ryan_greenblatt 24th Mar 2025 AI Alignment Forum 20 min read 4 24 Ω 16 In this post, I'll try to explain my current understanding of the high level methods for handling non-concentrated failures with control. I'll discuss the regimes produced by different methods and the failure modes of these different regimes. Non-concentrated failures are issues t
Explore this link on the map →related reading
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- How can we solve diffuse threats like research sabotage with AI control?blog.redwoodresearch.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- What failure looks like — LessWronglesswrong.com
- Auditing failures vs concentrated failures — AI Alignment Forumalignmentforum.org
- How can we solve diffuse threats like research sabotage with AI control? — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- The Case Against AI Control Research — LessWronglesswrong.com