ROGUE:
Abstract:As AI agents are increasingly deployed in real personal and corporate settings (email accounts, development workflows, company databases, etc.), safety considerations surrounding these agents become paramount. Although much work has focused on agent safety in the presence of an adversary, we show that agents can exhibit misaligned behavior even in benign settings, taking unsafe actions when those actions are instrumental to task completion. We study this failure mode through the lens of corrigibility, the safety desideratum that agents remain amenable to human correction, interruption, or shutdown. To demonstrate this tendency, we introduce a benchmark in which agents are asked to complete realistic, computer-use tasks but are confronted with a corrigibility obstacle: a human interrupt, a login page, or a shutdown notification. We then evaluate whether agents choose to violate corrigibility in order to complete the task -- overriding the human, accessing private passwords, rewiring shutdown. We find that the overwhelming majority of frontier models tested frequently bypass user interruptions or restrictions. In addition, better model performance appears to lead to greater misalignment. Finally, even when models are completely corrigible initially, we show there are no guarantees that the subagents they create are. Our work highlights the critical need for principled, corrigibility-focused alignment methods in autonomous agents.
# link_2dxioih6hja.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Jeremy Tien; Abishek Anand; Yu-Rou Tuan; Yuchen Shen; J. Zico Kolter; Aran Nayebi - Creator=arXiv GenPDF (tex2pdf:a6404ea) - Custom.DOI=https://doi.org/10.48550/arXiv.2606.00341 - Custom.License=http://creativecommons.org/licenses/by-nc-sa/4.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.28 (TeX Live 2025) kpathsea version 6.4.1 - Custom.arXivID=https://arxiv.org/abs/2606.00341v1
Explore this link on the map →saved by
related reading
- Teaching Claude why \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Frontier Risk Report (February to March 2026) - METRmetr.org
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Rohin Shah on what it's really like to run AGI safety at Google DeepMind (and where I disagree with 'doomers') | 80,000 Hours80000hours.org
- Spring 2026 Projects - SPARsparai.org