✳flâneur — a map of the web's best reading
Secretly Loyal AIs: Threat Vectors and Mitigation Strategies — LessWrong
lesswrong.com · 8,081 words · saved by 1 readers
Risks from secretly loyal AIs
x Secretly Loyal AIs: Threat Vectors and Mitigation Strategies — LessWrong AI Frontpage 8 Secretly Loyal AIs: Threat Vectors and Mitigation Strategies by Dave Banerjee 31st Oct 2025 Linkpost for substack.com 23 min read 0 8 Special thanks to my supervisor Girish Sastry for his guidance and support throughout this project. I am also grateful to Alan Chan, Houlton McGuinn, Connor Aidan Stewart Hunter, Rose Hadshar, Tom Davidson, and Cody Rushing for their valuable feedback on both the writing and conceptual development of this work. This post draws heavily upon Forethought’s report on AI-Enabled
Explore this link on the map →saved by
related reading
- secret-loyalties-whitepaper.pdfformationresearch.com
- AI Deterrence by Betrayalaibetrayal.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- AI-Enabled Coups: How a Small Group Could Use AI to Seize Powerforethought.org
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Off Target | CNAScnas.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Spring 2026 Projects - SPARsparai.org
- AI Integrity: Defending Against Backdoors and Secret Loyalties - Institute for AI Policy and Strategyiaps.ai
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org