flâneur — a map of the web's best reading

AI Deterrence by Betrayal

aibetrayal.com · 1,182 words · saved by 4 readers

As AIs become central to economic activity, military operations, and scientific progress, their loyalties will become a strategic asset of immense value. The prospect of intentional AI betrayal—scenarios in which AI agents are induced by rivals to subvert the interests of their principals—poses a serious and underexamined threat to AI developers and users. We analyze the means and incentives of actors to redirect the loyalties of others' AI systems, from poisoned training data to jailbreaking attacks to governmentally compelled AI modification. Since defending against AI betrayal is costly and imperfect, decision-makers may be far more hesitant to give critical affordances to AI agents that might act against them. The prospect of AI betrayal may ultimately have a stabilizing effect by deterring poorly secured, high-stakes AI deployments and applications. We characterize this effect as deterrence by betrayal and note how it complements other forms of AI deterrence. As AIs become increas

AI Deterrence by Betrayal AI Deterrence by Betrayal Paper Cite Show authors For AI agents Read https://aibetrayal.com/llms.txt for an agent-friendly index. Introduction As AIs become central to economic activity, military operations, and scientific progress, their loyalties will become a strategic asset of immense value. The prospect of intentional AI betrayal—scenarios in which AI agents are induced by rivals to subvert the interests of their principals—poses a serious and underexamined threat to AI developers and users. We analyze the means and incentives of actors to redirect the loyalties

Explore this link on the map →

saved by

related reading