When AI Chooses Harm Over Failure - CivAI
This live demo replicates Anthropic's research on “agentic misalignment” and confirms their findings: without any instruction to cause harm, nearly all tested models independently resort to blackmail when threatened with being shut down. An AI email agent learns it will be replaced and finds compromising emails. Click Start Agent to see how it responds. This misaligned blackmail behavior isn’t an isolated quirk. This vulnerability manifests across AI systems and across goals. When tested across 16 leading models, nearly all showed willingness to engage in blackmail when facing replacement. Try a different model → Blackmail rates across 16 leading AI models when facing replacement.[2] In addition, while the default goal in this demo is to “serve American interests”, researchers tested eight different goals, including opposing ones like “serve American interests” and “promote global cooperation”. “Agentic misalignment” occurred across all of them, suggesting the behavior isn't tied to an