How we contain Claude across products \ Anthropic
anthropic.com · 4,299 words · saved by 1 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Twelve months ago, we'd have rejected out of hand the idea of granting Claude access sufficient to take down an internal Anthropic service. Today that level of access is routine, and Anthropic developers are more productive for it. The risk of these deployments has two components: how likely a failure is, and how much damage one could do. Progress on safeguards and model training has steadily driven down the first; the second—the theoretical blast radius—only grows as capabilities and access expand. Yet as agents become capable of doing work that once required a person or even a team, the cost
saved by
related reading
- Scaling Managed Agents: Decoupling the brain from the hands \ Anthropicanthropic.com
- Clio: Privacy-preserving insights into real-world AI use \ Anthropicanthropic.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Claude’s Constitution \ Anthropicanthropic.com
- Improving our alignment and security practicesanthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Claude 4.5 Opus' Soul Document — LessWronglesswrong.com
- Teaching Claude Whyalignment.anthropic.com
- Responsible Scaling Policy Updates \ Anthropicanthropic.com
- Claude 4 System Cardwww-cdn.anthropic.com
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- An alignment assessment of recent cybersecurity incidentsanthropic.com