flâneur — a map of the web's best reading

Coup probes: Catching catastrophes with probes trained off-policy — LessWrong

lesswrong.com · 5,286 words · saved by 1 readers

Thanks to Ryan Greenblatt for the original description of the ideas and hopes, and to Erik Jenner, Jean-Stanislas Denain, Nate Thomas, Ryan Greenblatt, and Buck Shlegeris for useful discussions and feedback on the draft of this post. Here’s a simple safety intervention: train a probe on your AI to tell you whether it’s currently trying to do a do a catastrophic action. We call these coup probes. This can be used to cheaply flag potential takeover attempts by early transformative AIs, as a complement to: How do you get the training data for coup probes? In this post, we suggest having humans generate synthetic examples of coup attempts (e.g. writing and exploiting a serious backdoor). Even if the model recognizes that the coup attempts are off-policy (i.e. it would have put very low probability on them), the probe may still partially capture the model’s concept of a coup attempt. It is unclear how successfully such probes would flag malign behavior, but they seem like a useful additiona

x Coup probes: Catching catastrophes with probes trained off-policy — LessWrong AI Control Redwood Research AI Frontpage 95 Coup probes: Catching catastrophes with probes trained off-policy by Fabien Roger 17th Nov 2023 AI Alignment Forum 13 min read 9 95 Ω 55 Thanks to Ryan Greenblatt for the original description of the ideas and hopes, and to Erik Jenner, Jean-Stanislas Denain, Nate Thomas, Ryan Greenblatt, and Buck Shlegeris for useful discussions and feedback on the draft of this post. Here’s a simple safety intervention: train a probe on your AI to tell you whether it’s currently trying t

Explore this link on the map →

related reading