Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols? — AI Alignment Forum
We recently released Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?, a major update to our previous paper/blogpost, evaluating a broader range of models (e.g. helpful-only Claude 3.5 Sonnet) in more diverse and realistic settings (e.g. untrusted monitoring). An AI control protocol is a plan for usefully deploying AI systems that prevents an AI from intentionally causing some unacceptable outcome. Previous work evaluated protocols by subverting them using an AI following a human-written strategy. This paper investigates how well AI systems can generate and act on their own strategies for subverting control protocols whilst operating statelessly (i.e., without shared memory between contexts). To do this, an AI system may need to reliably generate effective strategies in each context, take actions with well-calibrated probabilities, and coordinate plans with other instances of itself without communicating. We develop Subversion Strategy
x Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols? — AI Alignment Forum AI Evaluations AI Frontpage 24 Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols? by Alex Mallen , Charlie Griffin , Buck 24th Mar 2025 10 min read 0 24 We recently released Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols? , a major update to our previous paper / blogpost , evaluating a broader range of models (e.g. helpful-only Claude 3.5 Sonnet) in more diverse and realistic
Explore this link on the map →related reading
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- 2312.06942arxiv.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionar5iv.labs.arxiv.org
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWronglesswrong.com
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org