Why imperfect adversarial robustness doesn't doom AI control
(thanks to Alex Mallen, Cody Rushing, Zach Stein-Perlman, Hoagy Cunningham, Vlad Mikulik, and Fabien Roger for comments) Sometimes I hear people argue against AI control as follows: if your control measures rely on getting good judgments from "trusted" AI models, you're doomed because AI models aren't adversarially robust. I agree that better adversarial robustness would definitely help with control, but I think this argument misses several important points; my overall take is that current levels of adversarial robustness don’t come close to dooming AI control. The first and most important point is that adversarial robustness is normally researched in a setting that’s much easier for the adversary. In most discussions, if an adversary can ever get a model to do something bad—even after many attempts and with a low success rate—this counts as a failure. But in real control settings, an AI that tries and fails to jailbreak might get caught, which is bad for its chances of success. In add
Why imperfect adversarial robustness doesn't doom AI control There are crucial disanalogies between preventing jailbreaks and preventing misalignment-induced catastrophes. Buck Shlegeris Nov 18, 2024 5 3 Share (thanks to Alex Mallen, Cody Rushing, Zach Stein-Perlman, Hoagy Cunningham, Vlad Mikulik, and Fabien Roger for comments) Sometimes I hear people argue against AI control as follows: if your control measures rely on getting good judgments from "trusted" AI models, you're doomed because AI models aren't adversarially robust. I agree that better adversarial robustness would definitely help
Explore this link on the map →related reading
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- gdm-ai-control-roadmap.pdfstorage.googleapis.com
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- 2312.06942arxiv.org
- Introduction to AI Control - by Sarah - BlueDot Impactblog.bluedot.org