flâneur

Win/continue/lose scenarios and execute/replace/audit protocols

blog.redwoodresearch.org · 1,977 words · saved by 1 readers

In this post, I’ll make a technical point that comes up when thinking about risks from scheming AIs from a control perspective.

In this post, I’ll make a technical point that comes up when thinking about risks from scheming AIs from a control perspective. In brief: Consider a deployment of an AI in a setting where it’s going to be given a sequence of tasks, and you’re worried about a safety failure that can happen suddenly. Every time the AI tries to attack, one of three things happens: we win, we lose, or the deployment continues (and the AI probably eventually attempts to attack again). So there’s two importantly different notions of attacks “failing”: we can catch the attack (in which case we win), or the attack…

saved by

related reading