Thoughts on Claude Fable's silent safeguards — LessWrong
lesswrong.com · 5,521 words · saved by 1 readers
[Update (June 11, 2026): Anthropic has since "un-silenced" the new safeguards (source).] …
x Thoughts on Claude Fable's silent safeguards — LessWrong AI Personal Blog 51 Thoughts on Claude Fable's silent safeguards by Andy Arditi 10th Jun 2026 12 min read 20 51 [Update (June 11, 2026): Anthropic has since "un-silenced" the new safeguards ( source ).] [Thanks to Julian Minder for helpful discussion and review.] Claude Fable 5 and its new safeguards Yesterday, Anthropic publicly released Claude Fable 5. Fable 5 is a Mythos-class model – a model class above Opus, Anthropic's previous premium tier – and, as assessed by multiple benchmarks, it is the most capable model to date. Due to th
saved by
related reading
- Pre-deployment auditing can catch an overt saboteuralignment.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Teaching Claude Whyalignment.anthropic.com
- Introducing Claude Fable 5.1 and Claude Mythos 5.1anthropic.com
- Claude Fable 5 and Claude Mythos 5 \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Claude’s Constitution \ Anthropicanthropic.com
- Anthropic’s Safety Superpower – Stratechery by Ben Thompsonstratechery.com
- Responsible Scaling Policy Updates \ Anthropicanthropic.com
- Redeploying Claude Fable 5 \ Anthropicanthropic.com