Teaching Claude Why
Last year, we released a case study on agentic misalignment. This research showed that AI models across the industry sometimes took egregiously misaligned actions when placed in (fictional) ethical dilemmas—for example, blackmailing engineers to avoid being shut down. At the time of this research, Claude 4 was Anthropic's frontier model family. It was also the first model family for which we ran a live alignment assessment during training, and agentic misalignment was one of several issues that surfaced (others include increased susceptibility to jailbreaks and harmful system prompting). So after Claude 4, it was clear we needed to improve our safety training. However, it was not initially clear what was driving the failures, or which kinds of interventions would generalize beyond the specific scenarios we had caught. Since then, we have significantly updated our safety training using methods discussed in this post as well as a number of more prosaic improvements to our training data,
Teaching Claude Why Alignment Science Blog Teaching Claude Why Jonathan Kutasov * , Adam Jermyn May 8, 2026 Julius Steen, Minh Le, Samuel R. Bowman, Samuel Marks, Jan Leike, Amanda Askell, Chris Olah Evan Hubinger, Sara Price * Correspondence to jonk@anthropic.com Introduction Last year, we released a case study on agentic misalignment . This research showed that AI models across the industry sometimes took egregiously misaligned actions when placed in (fictional) ethical dilemmas—for example, blackmailing engineers to avoid being shut down. At the time of this research, Claude 4 was Anthropic
saved by
related reading
- Teaching Claude why \ Anthropicanthropic.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- An alignment assessment of recent cybersecurity incidentsanthropic.com
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Model Spec Midtraining: Improving How Alignment Training Generalizesalignment.anthropic.com
- Claude Opus 4.5: Model Card, Alignment and Safetythezvi.substack.com
- Claude’s Character \ Anthropicanthropic.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com