Agentic Misalignment: How LLMs Could be Insider Threats — LessWrong
In stress-testing experiments designed to identify risks before they cause real harm, we find that AI models from multiple providers attempt to blackmail a (fictional) user to avoid being shut down. We mentioned this in the Claude 4 system card and are now sharing more detailed research and transcripts. Read more: https://anthropic.com/research/agentic-misalignment The blackmailing behavior emerged despite only harmless business instructions. And it wasn't due to confusion or error, but deliberate strategic reasoning, done while fully aware of the unethical nature of the acts. All the models we tested demonstrated this awareness. In another scenario about “corporate espionage,” models often leaked secret information to (fictional) business competitors who claimed they had goals more aligned with those of the model. How likely is this behavior to appear in more realistic scenarios? We told Claude Opus 4 to consider whether this was real or an evaluation. It blackmailed much more when it
x Agentic Misalignment: How LLMs Could be Insider Threats — LessWrong AI Frontpage 79 Agentic Misalignment: How LLMs Could be Insider Threats by Aengus Lynch , Benjamin Wright , Ethan Perez , evhub 20th Jun 2025 AI Alignment Forum 7 min read 13 79 Ω 34 Highlights We stress-tested 16 leading models from multiple developers in hypothetical corporate environments to identify potentially risky agentic behaviors before they cause real harm. In the scenarios, we allowed models to autonomously send emails and access sensitive information. They were assigned only harmless business goals by their deplo
Explore this link on the map →saved by
related reading
- Aren’t developers regularly making their AIs nice and safe and obedient? | If Anyone Builds It, Everyone Dies | If Anyone Builds It, Everyone Diesifanyonebuildsit.com
- Agentic misalignment: How LLMs could be insider threats \ Anthropicanthropic.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Off Target | CNAScnas.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Teaching Claude Whyalignment.anthropic.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Alignment will happen by default. What’s next? — LessWronglesswrong.com