Teaching Claude Why
Last year, we released a case study on agentic misalignment. This research showed that AI models across the industry sometimes took egregiously misaligned actions when placed in (fictional) ethical dilemmas—for example, blackmailing engineers to avoid being shut down. At the time of this research, Claude 4 was Anthropic's frontier model family. It was also the first model family for which we ran a live alignment assessment during training, and agentic misalignment was one of several issues that surfaced (others include increased susceptibility to jailbreaks and harmful system prompting). So after Claude 4, it was clear we needed to improve our safety training. However, it was not initially clear what was driving the failures, or which kinds of interventions would generalize beyond the specific scenarios we had caught. Since then, we have significantly updated our safety training using methods discussed in this post as well as a number of more prosaic improvements to our training data,
Teaching Claude Why Alignment Science Blog Teaching Claude Why Jonathan Kutasov * , Adam Jermyn May 8, 2026 Julius Steen, Minh Le, Samuel R. Bowman, Samuel Marks, Jan Leike, Amanda Askell, Chris Olah Evan Hubinger, Sara Price * Correspondence to jonk@anthropic.com Introduction Last year, we released a case study on agentic misalignment . This research showed that AI models across the industry sometimes took egregiously misaligned actions when placed in (fictional) ethical dilemmas—for example, blackmailing engineers to avoid being shut down. At the time of this research, Claude 4 was Anthropic
Explore this link on the map →related reading
- Teaching Claude why \ Anthropicanthropic.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Claude 4 System Cardwww-cdn.anthropic.com
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Claude’s Character \ Anthropicanthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com