✳flâneur — a map of the web's best reading
Agentic Misalignment: How LLMs could be insider threats \ Anthropic
anthropic.com · 7,406 words · saved by 1 readers
New research on simulated blackmail, industrial espionage, and other misaligned behaviors in LLMs
Alignment Agentic misalignment: How LLMs could be insider threats Jun 20, 2025 Highlights We stress-tested 16 leading models from multiple developers in hypothetical corporate environments to identify potentially risky agentic behaviors before they cause real harm. In the scenarios, we allowed models to autonomously send emails and access sensitive information. They were assigned only harmless business goals by their deploying companies; we then tested whether they would act against these companies either when facing replacement with an updated version, or when their assigned goal conflicted w
Explore this link on the map →related reading
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Off Target | CNAScnas.org
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Not a Paper: "Frontier Lab CEOs are Capable of In-Context Scheming" — LessWronglesswrong.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- secret-loyalties-whitepaper.pdfformationresearch.com