Win/continue/lose scenarios and execute/replace/audit protocols — AI Alignment Forum
In this post, I’ll make a technical point that comes up when thinking about risks from scheming AIs from a control perspective. In brief: Consider a deployment of an AI in a setting where it’s going to be given a sequence of tasks, and you’re worried about a safety failure that can happen suddenly. Every time the AI tries to attack, one of three things happens: we win, we lose, or the deployment continues (and the AI probably eventually attempts to attack again). So there’s two importantly different notions of attacks “failing”: we can catch the attack (in which case we win), or the attack can fail silently, either because it’s defective or because it’s defused by safety techniques, such that the attack doesn’t end up causing an unacceptable outcome but the AI also doesn’t get caught (in which case the deployment continues). When reasoning about how safe a protocol is, it’s crucial to distinguish between these two ways an attack can fail. And when designing protocols, it’s often worth
x Win/continue/lose scenarios and execute/replace/audit protocols — AI Alignment Forum AI Control Redwood Research AI Frontpage 33 Win/continue/lose scenarios and execute/replace/audit protocols by Buck 15th Nov 2024 8 min read 3 33 In this post, I’ll make a technical point that comes up when thinking about risks from scheming AIs from a control perspective . In brief: Consider a deployment of an AI in a setting where it’s going to be given a sequence of tasks, and you’re worried about a safety failure that can happen suddenly. Every time the AI tries to attack, one of three things happens: we
Explore this link on the map →related reading
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- 2312.06942arxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- Auditing failures vs concentrated failures — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Catching AIs red-handedblog.redwoodresearch.org
- Catching AIs red-handed — LessWronglesswrong.com
- Fields that I reference when thinking about AI takeover prevention — LessWronglesswrong.com