AI Deterrence by Betrayal
As AIs become central to economic activity, military operations, and scientific progress, their loyalties will become a strategic asset of immense value. The prospect of intentional AI betrayal—scenarios in which AI agents are induced by rivals to subvert the interests of their principals—poses a serious and underexamined threat to AI developers and users. We analyze the means and incentives of actors to redirect the loyalties of others' AI systems, from poisoned training data to jailbreaking attacks to governmentally compelled AI modification. Since defending against AI betrayal is costly and imperfect, decision-makers may be far more hesitant to give critical affordances to AI agents that might act against them. The prospect of AI betrayal may ultimately have a stabilizing effect by deterring poorly secured, high-stakes AI deployments and applications. We characterize this effect as deterrence by betrayal and note how it complements other forms of AI deterrence. As AIs become increas
AI Deterrence by Betrayal AI Deterrence by Betrayal Paper Cite Show authors For AI agents Read https://aibetrayal.com/llms.txt for an agent-friendly index. Introduction As AIs become central to economic activity, military operations, and scientific progress, their loyalties will become a strategic asset of immense value. The prospect of intentional AI betrayal—scenarios in which AI agents are induced by rivals to subvert the interests of their principals—poses a serious and underexamined threat to AI developers and users. We analyze the means and incentives of actors to redirect the loyalties
Explore this link on the map →saved by
related reading
- Secretly Loyal AIs: Threat Vectors and Mitigation Strategies — LessWronglesswrong.com
- The Three Filters: Why Almost Every Plan to Survive ASI Fails Miserably — LessWronglesswrong.com
- Strategic Visionsstatic1.squarespace.com
- Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence Strategynationalsecurity.ai
- The Moonshot - by Anton Leicht - Threading the Needlewriting.antonleicht.me
- AI-Enabled Coups: How a Small Group Could Use AI to Seize Powerforethought.org
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- AI 2027ai-2027.com
- secret-loyalties-whitepaper.pdfformationresearch.com
- Executive Summary — Chapter 1 of Superintelligence Strategynationalsecurity.ai
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Seeking Stability in the Competition for AI Advantage | RANDrand.org