Coup probes: Catching catastrophes with probes trained off-policy — LessWrong
Thanks to Ryan Greenblatt for the original description of the ideas and hopes, and to Erik Jenner, Jean-Stanislas Denain, Nate Thomas, Ryan Greenblatt, and Buck Shlegeris for useful discussions and feedback on the draft of this post. Here’s a simple safety intervention: train a probe on your AI to tell you whether it’s currently trying to do a do a catastrophic action. We call these coup probes. This can be used to cheaply flag potential takeover attempts by early transformative AIs, as a complement to: How do you get the training data for coup probes? In this post, we suggest having humans generate synthetic examples of coup attempts (e.g. writing and exploiting a serious backdoor). Even if the model recognizes that the coup attempts are off-policy (i.e. it would have put very low probability on them), the probe may still partially capture the model’s concept of a coup attempt. It is unclear how successfully such probes would flag malign behavior, but they seem like a useful additiona
x Coup probes: Catching catastrophes with probes trained off-policy — LessWrong AI Control Redwood Research AI Frontpage 95 Coup probes: Catching catastrophes with probes trained off-policy by Fabien Roger 17th Nov 2023 AI Alignment Forum 13 min read 9 95 Ω 55 Thanks to Ryan Greenblatt for the original description of the ideas and hopes, and to Erik Jenner, Jean-Stanislas Denain, Nate Thomas, Ryan Greenblatt, and Buck Shlegeris for useful discussions and feedback on the draft of this post. Here’s a simple safety intervention: train a probe on your AI to tell you whether it’s currently trying t
Explore this link on the map →related reading
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- 2312.06942arxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- [2601.11516] Building Production-Ready Probes For Geminiarxiv.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- How well do truth probes generalise? — LessWronglesswrong.com
- Catching AIs red-handedblog.redwoodresearch.org
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- Catching AIs red-handed — LessWronglesswrong.com