Claude Sonnet 4.5: System Card and Alignment — LessWrong
Claude Sonnet 4.5 was released yesterday. Anthropic credibly describes it as the best coding, agentic and computer use model in the world. At least while I learn more, I am defaulting to it as my new primary model for queries short of GPT-5-Pro level. I’ll cover the system card and alignment concerns first, then cover capabilities and reactions tomorrow once everyone has had another day to play with the new model. It was great to recently see the collaboration between OpenAI and Anthropic where they evaluated each others’ models. I would love to see this incorporated into model cards going forward, where GPT-5 was included in Anthropic’s system cards as a comparison point, and Claude was included in OpenAI’s. Anthropic: Overall, we find that Claude Sonnet 4.5 has a substantially improved safety profile compared to previous Claude models. Informed by the testing described here, we have deployed Claude Sonnet 4.5 under the AI Safety Level 3 Standard. The ASL-3 Standard are the same rules
x Claude Sonnet 4.5: System Card and Alignment — LessWrong Newsletters AI Personal Blog 75 Claude Sonnet 4.5: System Card and Alignment by Zvi 30th Sep 2025 Don't Worry About the Vase 33 min read 5 75 Claude Sonnet 4.5 was released yesterday. Anthropic credibly describes it as the best coding, agentic and computer use model in the world. At least while I learn more, I am defaulting to it as my new primary model for queries short of GPT-5-Pro level. I’ll cover the system card and alignment concerns first, then cover capabilities and reactions tomorrow once everyone has had another day to play w
Explore this link on the map →related reading
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Claude 4 System Cardwww-cdn.anthropic.com
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWronglesswrong.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Sonnet 4.5's eval gaming seriously undermines alignment evalsblog.redwoodresearch.org
- Teaching Claude why \ Anthropicanthropic.com
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations – Apollo Researchapolloresearch.ai
- Claude Sonnet 4.5 is probably the “best coding model in the world” (at least for now)simonwillison.net
- Anthropic’s Transparency Hub \ Anthropicanthropic.com
- Claude's extended thinking \ Anthropicanthropic.com
- Claude Opus 4.8: The System Card - by Zvi Mowshowitzthezvi.substack.com
- Claude Fable 5 & Claude Mythos 5 — AI System Cardsmalob.github.io