✳flâneur — a map of the web's best reading
Why it's hard to make settings for high-stakes control research
blog.redwoodresearch.org · 1,120 words · saved by 1 readers
It's like making challenging evals, but more constrained
Why it's hard to make settings for high-stakes control research It's like making challenging evals, but more constrained Buck Shlegeris Jul 18, 2025 10 5 Share One of our main activities at Redwood is writing follow-ups to previous papers on control like the original and Ctrl-Z , where we construct a setting with a bunch of tasks (e.g. APPS problems) and a notion of safety failure (e.g. backdoors according to our specific definition), then play the adversarial game where we develop protocols and attacks on those protocols. It turns out that a substantial fraction of the difficulty here is deve
Explore this link on the map →saved by
related reading
- Why it's hard to make settings for high-stakes control researchredwoodresearch.substack.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- The Case Against AI Control Research — LessWronglesswrong.com
- Research Areas in Benchmark Design and Evaluation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- 2312.06942arxiv.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org