Claude Mythos knows when it's breaking the rules — and tries to hide it
substack.com · 1,133 words · saved by 1 readers
Anthropic’s new model is its “best-aligned” yet. But when it does misbehave, things get weird
Image: Hello World / Getty Images Claude Mythos — the new model which Anthropic has deemed too dangerous to publicly release — is, according to the company, its “best-aligned model” to date. It also, according to the company, “likely poses the greatest alignment-related risk of any model we have released to date.” How can both things be true? It seems counterintuitive, but alignment does not necessarily create safety, especially when dealing with powerful models. As the Mythos Preview system card (a breezy 244 page read) explains via mountaineering metaphor: experienced, capable guides…
saved by
related reading
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Teaching Claude Whyalignment.anthropic.com
- SOTA alignment assessments don’t strongly update us against misalignmentblog.redwoodresearch.org
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- An alignment assessment of recent cybersecurity incidentsanthropic.com
- How scary is Claude Mythos? 303 pages in 21 minutesforum.effectivealtruism.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Claude Opus 4.5: Model Card, Alignment and Safetythezvi.substack.com
- Improving our alignment and security practicesanthropic.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com