flâneur — a map of the web's best reading

Agentic Misalignment: How LLMs Could be Insider Threats — LessWrong

lesswrong.com · 3,379 words · saved by 1 readers

In stress-testing experiments designed to identify risks before they cause real harm, we find that AI models from multiple providers attempt to blackmail a (fictional) user to avoid being shut down. We mentioned this in the Claude 4 system card and are now sharing more detailed research and transcripts. Read more: https://anthropic.com/research/agentic-misalignment The blackmailing behavior emerged despite only harmless business instructions. And it wasn't due to confusion or error, but deliberate strategic reasoning, done while fully aware of the unethical nature of the acts. All the models we tested demonstrated this awareness. In another scenario about “corporate espionage,” models often leaked secret information to (fictional) business competitors who claimed they had goals more aligned with those of the model. How likely is this behavior to appear in more realistic scenarios? We told Claude Opus 4 to consider whether this was real or an evaluation. It blackmailed much more when it

x Agentic Misalignment: How LLMs Could be Insider Threats — LessWrong AI Frontpage 79 Agentic Misalignment: How LLMs Could be Insider Threats by Aengus Lynch , Benjamin Wright , Ethan Perez , evhub 20th Jun 2025 AI Alignment Forum 7 min read 13 79 Ω 34 Highlights We stress-tested 16 leading models from multiple developers in hypothetical corporate environments to identify potentially risky agentic behaviors before they cause real harm. In the scenarios, we allowed models to autonomously send emails and access sensitive information. They were assigned only harmless business goals by their deplo

Explore this link on the map →

saved by

related reading