secret-loyalties-whitepaper.pdf
formationresearch.com · 13,664 words · saved by 3 readers
N/A
# link_1bh4vexchsn.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=true - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - CreationDate=D:20260513153137Z - Creator=LaTeX with hyperref - ModDate=D:20260513153137Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.27 (TeX Live 2025) kpathsea version 6.4.1 - Producer=pdfTeX-1.40.27 - Trapped=False ## Contents ### Page 1 AIs with Secret Loyalties are a Serious but Addressable ThreatJoe Kwon∗ Alfie LamertonFormation ResearchAndrew DraganovArcadia ImpactDave BanerjeeIA
saved by
related reading
- Secretly Loyal AIs: Threat Vectors and Mitigation Strategies — LessWronglesswrong.com
- Thoughts on Claude Fable's silent safeguards — LessWronglesswrong.com
- Notes on Inference Integritynewsletter.forethought.org
- We need 3rd party Training-Run Assessments — LessWronglesswrong.com
- Pre-deployment auditing can catch an overt saboteuralignment.anthropic.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- [2606.04071] Covert Influence Between Language Modelsarxiv.org
- Off Target | CNAScnas.org
- Auditing language models for hidden objectives — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI Integrity: Defending Against Backdoors and Secret Loyalties - Institute for AI Policy and Strategyiaps.ai