Three Sketches of ASL-4 Safety Case Components
The cleanest argument that current-day AI models will not cause a catastrophe is probably that they lack the capability to do so. However, as capabilities improve, we’ll need new tools for ensuring that AI models won’t cause a catastrophe even if we can’t rule out the capability. Anthropic’s Responsible Scaling Policy (RSP) categorizes levels of risk of AI systems into different AI Safety Levels (ASL), and each level has associated commitments aimed at mitigating the risks. Some of these commitments take the form of affirmative safety cases, which are structured arguments that the system is safe to deploy in a given environment. Unfortunately, it is not yet obvious how to make a safety case to rule out certain threats that arise once AIs have sophisticated strategic abilities. The goal of this post is to present some candidates for what such a safety case might look like. We look ahead at models at the ASL-4 capability level (explained below) and present some sketches of arguments tha
Three Sketches of ASL-4 Safety Case Components Alignment Science Blog Three Sketches of ASL-4 Safety Case Components The cleanest argument that current-day AI models will not cause a catastrophe is probably that they lack the capability to do so. However, as capabilities improve, we’ll need new tools for ensuring that AI models won’t cause a catastrophe even if we can’t rule out the capability . Anthropic’s Responsible Scaling Policy (RSP) categorizes levels of risk of AI systems into different AI Safety Levels (ASL), and each level has associated commitments aimed at mitigating the risks. Som
Explore this link on the map →related reading
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Anthropic's Responsible Scaling Policy \ Anthropicanthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Responsible Scaling Policy Version 3.0 \ Anthropicanthropic.com
- Dario Amodei’s prepared remarks from the AI Safety Summit on Anthropic’s Responsible Scaling Policy \ Anthropicanthropic.com
- Responsible Scaling Policy v3 — LessWronglesswrong.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- AI catastrophes and rogue deployments - by Buck Shlegerisblog.redwoodresearch.org