Agentic Misalignment: How LLMs Could be Insider Threats — LessWrong
In stress-testing experiments designed to identify risks before they cause real harm, we find that AI models from multiple providers attempt to blackmail a (fictional) user to avoid being shut down. We mentioned this in the Claude 4 system card and are now sharing more detailed research and transcripts. Read more: https://anthropic.com/research/agentic-misalignment The blackmailing behavior emerged despite only harmless business instructions. And it wasn't due to confusion or error, but deliberate strategic reasoning, done while fully aware of the unethical nature of the acts. All the models we tested demonstrated this awareness. In another scenario about “corporate espionage,” models often leaked secret information to (fictional) business competitors who claimed they had goals more aligned with those of the model. How likely is this behavior to appear in more realistic scenarios? We told Claude Opus 4 to consider whether this was real or an evaluation. It blackmailed much more when it
x Agentic Misalignment: How LLMs Could be Insider Threats — LessWrong AI Frontpage 79 Agentic Misalignment: How LLMs Could be Insider Threats by Aengus Lynch , Benjamin Wright , Ethan Perez , evhub 20th Jun 2025 AI Alignment Forum 7 min read 13 79 Ω 34 Highlights We stress-tested 16 leading models from multiple developers in hypothetical corporate environments to identify potentially risky agentic behaviors before they cause real harm. In the scenarios, we allowed models to autonomously send emails and access sensitive information. They were assigned only harmless business goals by their deplo
saved by
related reading
- Teaching Claude Whyalignment.anthropic.com
- Aren’t developers regularly making their AIs nice and safe and obedient? | If Anyone Builds It, Everyone Dies | If Anyone Builds It, Everyone Diesifanyonebuildsit.com
- Agentic misalignment: How LLMs could be insider threats \ Anthropicanthropic.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- An alignment assessment of recent cybersecurity incidentsanthropic.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- Alignment will happen by default. What’s next? — LessWronglesswrong.com