Why it's hard to make settings for high-stakes control research
One of our main activities at Redwood is writing follow-ups to previous papers on control like the original and Ctrl-Z, where we construct a setting with a bunch of tasks (e.g. APPS problems) and a notion of safety failure (e.g. backdoors according to our specific definition), then play the adversarial game where we develop protocols and attacks on those protocols. It turns out that a substantial fraction of the difficulty here is developing the settings, rather than developing the protocols and attacks. I want to explain why this is the case. In order to do high-stakes control research, you need a dataset of tasks with the following properties: You can check whether the main task was completed successfully. There's some notion of safety failure that we consider unacceptable; we can check whether a safety failure occurred. (We sometimes call these safety failures "side tasks".) Here's why it's challenging to develop settings that work. Focusing entirely on the main tasks to start with:
Why it's hard to make settings for high-stakes control research It's like making challenging evals, but more constrained Buck Shlegeris Jul 18, 2025 11 5 Share One of our main activities at Redwood is writing follow-ups to previous papers on control like the original and Ctrl-Z , where we construct a setting with a bunch of tasks (e.g. APPS problems) and a notion of safety failure (e.g. backdoors according to our specific definition), then play the adversarial game where we develop protocols and attacks on those protocols. It turns out that a substantial fraction of the difficulty here is deve
Explore this link on the map →related reading
- Why it's hard to make settings for high-stakes control researchblog.redwoodresearch.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Research Areas in Benchmark Design and Evaluation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- The Case Against AI Control Research — LessWronglesswrong.com
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- 2312.06942arxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org