Ctrl-Z: Controlling AI Agents via Resampling — LessWrong
We have released a new paper, Ctrl-Z: Controlling AI Agents via Resampling. This is the largest and most intricate study of control techniques to date: that is, techniques that aim to prevent catastrophic failures even if egregiously misaligned AIs attempt to subvert the techniques. We extend control protocols to a more realistic, multi-step setting, develop novel resample protocols that outperform all existing control protocols, and study many aspects of control techniques in more detail than prior work. In this blog post, we summarize the main takeaways and lessons learned from our work. Here’s the abstract of the paper: Control evaluations measure whether monitoring and security protocols for AI systems prevent intentionally subversive AI models from causing harm. Our work presents the first control evaluation performed in an agent environment. We construct BashBench, a dataset of 257 challenging multi-step system administration tasks, and evaluate whether various safety measures ca
x Ctrl-Z: Controlling AI Agents via Resampling — LessWrong AI Control MATS Program AI Frontpage 2025 Top Fifty: 14 % 128 Ctrl-Z: Controlling AI Agents via Resampling by Aryan Bhatt , Buck , Adam Kaufman , Tyler Tracy 16th Apr 2025 AI Alignment Forum 24 min read 0 128 Ω 56 We have released a new paper, Ctrl-Z: Controlling AI Agents via Resampling . This is the largest and most intricate study of control techniques to date: that is, techniques that aim to prevent catastrophic failures even if egregiously misaligned AIs attempt to subvert the techniques. We extend control protocols to a more real
Explore this link on the map →related reading
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- [2504.10374] Ctrl-Z: Controlling AI Agents via Resamplingarxiv.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- 2312.06942arxiv.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- Recent Redwood Research project proposals — AI Alignment Forumalignmentforum.org
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org