Win/continue/lose scenarios and execute/replace/audit protocols
blog.redwoodresearch.org · 1,977 words · saved by 1 readers
In this post, I’ll make a technical point that comes up when thinking about risks from scheming AIs from a control perspective.
In this post, I’ll make a technical point that comes up when thinking about risks from scheming AIs from a control perspective. In brief: Consider a deployment of an AI in a setting where it’s going to be given a sequence of tasks, and you’re worried about a safety failure that can happen suddenly. Every time the AI tries to attack, one of three things happens: we win, we lose, or the deployment continues (and the AI probably eventually attempts to attack again). So there’s two importantly different notions of attacks “failing”: we can catch the attack (in which case we win), or the attack…
saved by
related reading
- How can we solve diffuse threats like research sabotage with AI control?blog.redwoodresearch.org
- Win/continue/lose scenarios and execute/replace/audit protocols — AI Alignment Forumalignmentforum.org
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Reading Listblog.redwoodresearch.org
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- 2312.06942arxiv.org
- Catching AIs red-handedblog.redwoodresearch.org
- What failure looks like — LessWronglesswrong.com
- Catching AIs red-handed — LessWronglesswrong.com