flâneur — a map of the web's best reading

Claude Opus 4.5: Model Card, Alignment and Safety

thezvi.substack.com · saved by 2 readers

The contrast in model cards is stark. Google provided a brief overview of its tests for Gemini 3 Pro, with a lot of ‘we did this test, and we learned a lot from it, and we are not going to tell you the results.’ Anthropic gives us a 150 page book, including their capability assessments. This makes sense. Capability is directly relevant to safety, and also frontier capability safety tests often also credible indications of capability. Which still has several instances of ‘we did this test, and we learned a lot from it, and we are not going to tell you the results.’ Damn it. I get it, but damn it. Anthropic claims Opus 4.5 is the most aligned frontier model to date, although ‘with many subtleties.’ I agree with Anthropic’s assessment, especially for practical purposes right now. Claude is also miles ahead of other models on aspects of alignment that do not directly appear on a frontier safety assessment. In terms of surviving superintelligence, it’s still the scene from The Phantom Menac

Explore this link on the map →

saved by