Activating AI Safety Level 3 protections \ Anthropic
We have activated the AI Safety Level 3 (ASL-3) Deployment and Security Standards described in Anthropic’s Responsible Scaling Policy (RSP) in conjunction with launching Claude Opus 4. The ASL-3 Security Standard involves increased internal security measures that make it harder to steal model weights, while the corresponding Deployment Standard covers a narrowly targeted set of deployment measures designed to limit the risk of Claude being misused specifically for the development or acquisition of chemical, biological, radiological, and nuclear (CBRN) weapons. These measures should not lead Claude to refuse queries except on a very narrow set of topics. We are deploying Claude Opus 4 with our ASL-3 measures as a precautionary and provisional action. To be clear, we have not yet determined whether Claude Opus 4 has definitively passed the Capabilities Threshold that requires ASL-3 protections. Rather, due to continued improvements in CBRN-related knowledge and capabilities, we have dete
Policy Activating AI Safety Level 3 protections May 22, 2025 We have activated the AI Safety Level 3 (ASL-3) Deployment and Security Standards described in Anthropic’s Responsible Scaling Policy (RSP) in conjunction with launching Claude Opus 4. The ASL-3 Security Standard involves increased internal security measures that make it harder to steal model weights, while the corresponding Deployment Standard covers a narrowly targeted set of deployment measures designed to limit the risk of Claude being misused specifically for the development or acquisition of chemical, biological, radiological,
related reading
- Responsible Scaling Policy Updates \ Anthropicanthropic.com
- Anthropic’s Responsible Scaling Policy \ Anthropicanthropic.com
- Claude 4 System Cardwww-cdn.anthropic.com
- A Safe Path to Open Weights - Thinking Machines Labthinkingmachines.ai
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Anthropic's Responsible Scaling Policy \ Anthropicanthropic.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Claude Opus 4.5: Model Card, Alignment and Safetythezvi.substack.com
- Anthropic’s Transparency Hub \ Anthropicanthropic.com
- Redeploying Claude Fable 5 \ Anthropicanthropic.com
- How we contain Claude across products \ Anthropicanthropic.com
- Common Elements of Frontier AI Safety Policiesmetr.org