✳flâneur — a map of the web's best reading
secret-loyalties-whitepaper.pdf
formationresearch.com · 13,664 words · saved by 3 readers
N/A
# link_1bh4vexchsn.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=true - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - CreationDate=D:20260513153137Z - Creator=LaTeX with hyperref - ModDate=D:20260513153137Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.27 (TeX Live 2025) kpathsea version 6.4.1 - Producer=pdfTeX-1.40.27 - Trapped=False ## Contents ### Page 1 AIs with Secret Loyalties are a Serious but Addressable ThreatJoe Kwon∗ Alfie LamertonFormation ResearchAndrew DraganovArcadia ImpactDave BanerjeeIA
Explore this link on the map →saved by
related reading
- Secretly Loyal AIs: Threat Vectors and Mitigation Strategies — LessWronglesswrong.com
- Thoughts on Claude Fable's silent safeguards — LessWronglesswrong.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Off Target | CNAScnas.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Auditing language models for hidden objectives — LessWronglesswrong.com
- 2312.06942arxiv.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- AI Integrity: Defending Against Backdoors and Secret Loyalties - Institute for AI Policy and Strategyiaps.ai
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com