AI Deterrence by Betrayal
As AIs become central to economic activity, military operations, and scientific progress, their loyalties will become a strategic asset of immense value. The prospect of intentional AI betrayal—scenarios in which AI agents are induced by rivals to subvert the interests of their principals—poses a serious and underexamined threat to AI developers and users. We analyze the means and incentives of actors to redirect the loyalties of others' AI systems, from poisoned training data to jailbreaking attacks to governmentally compelled AI modification. Since defending against AI betrayal is costly and imperfect, decision-makers may be far more hesitant to give critical affordances to AI agents that might act against them. The prospect of AI betrayal may ultimately have a stabilizing effect by deterring poorly secured, high-stakes AI deployments and applications. We characterize this effect as deterrence by betrayal and note how it complements other forms of AI deterrence. As AIs become increas
AI Deterrence by Betrayal AI Deterrence by Betrayal Paper Cite Show authors For AI agents Read https://aibetrayal.com/llms.txt for an agent-friendly index. Introduction As AIs become central to economic activity, military operations, and scientific progress, their loyalties will become a strategic asset of immense value. The prospect of intentional AI betrayal—scenarios in which AI agents are induced by rivals to subvert the interests of their principals—poses a serious and underexamined threat to AI developers and users. We analyze the means and incentives of actors to redirect the loyalties
saved by
related reading
- Secretly Loyal AIs: Threat Vectors and Mitigation Strategies — LessWronglesswrong.com
- Strategic AI Sabotage: State Attacks on Advanced Systems' Developmenttwmstone.com
- The Three Filters: Why Almost Every Plan to Survive ASI Fails Miserably — LessWronglesswrong.com
- Strategic Visionsstatic1.squarespace.com
- Deterrence with Mutual Assured AI Malfunction (MAIM) — Chapter 4 of Superintelligence Strategynationalsecurity.ai
- AI Integrity: Defending Against Backdoors and Secret Loyalties - Institute for AI Policy and Strategyiaps.ai
- AI-Enabled Coups: How a Small Group Could Use AI to Seize Powerforethought.org
- The Moonshot - by Anton Leicht - Threading the Needlewriting.antonleicht.me
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Risk-Averse AIsforethought.org
- secret-loyalties-whitepaper.pdfformationresearch.com
- Executive Summary — Chapter 1 of Superintelligence Strategynationalsecurity.ai