Secretly Loyal AIs: Threat Vectors and Mitigation Strategies — LessWrong
lesswrong.com · 8,081 words · saved by 1 readers
Risks from secretly loyal AIs
x Secretly Loyal AIs: Threat Vectors and Mitigation Strategies — LessWrong AI Frontpage 8 Secretly Loyal AIs: Threat Vectors and Mitigation Strategies by Dave Banerjee 31st Oct 2025 Linkpost for substack.com 23 min read 0 8 Special thanks to my supervisor Girish Sastry for his guidance and support throughout this project. I am also grateful to Alan Chan, Houlton McGuinn, Connor Aidan Stewart Hunter, Rose Hadshar, Tom Davidson, and Cody Rushing for their valuable feedback on both the writing and conceptual development of this work. This post draws heavily upon Forethought’s report on AI-Enabled
saved by
related reading
- secret-loyalties-whitepaper.pdfformationresearch.com
- AI Integrity: Defending Against Backdoors and Secret Loyalties - Institute for AI Policy and Strategyiaps.ai
- AI Deterrence by Betrayalaibetrayal.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- AI-Enabled Coups: How a Small Group Could Use AI to Seize Powerforethought.org
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Off Target | CNAScnas.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- gdm-ai-control-roadmap.pdfstorage.googleapis.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Spring 2026 Projects - SPARsparai.org