✳flâneur — a map of the web's best reading
Teaching Claude why \ Anthropic
anthropic.com · 2,022 words · saved by 6 readers
New research on how we've reduced agentic misalignment
Alignment Teaching Claude why May 8, 2026 Last year, we released a case study on agentic misalignment . In experimental scenarios, we showed that AI models from many different developers sometimes took egregiously misaligned actions when they encountered (fictional) ethical dilemmas. For example, in one heavily discussed example, the models blackmailed engineers to avoid being shut down. When we first published this research, our most capable frontier models were from the Claude 4 family. This was also the first model family for which we ran a live alignment assessment during training; 1 agent
Explore this link on the map →saved by
related reading
- Teaching Claude Whyalignment.anthropic.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Claude’s Character \ Anthropicanthropic.com
- How far does alignment midtraining generalize?alignment.openai.com
- Claude is Now Alignment-Pretrained — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Claude 4 System Cardwww-cdn.anthropic.com