Sabotage evaluations for frontier models \ Anthropic
Any industry where there are potential harms needs evaluations. Nuclear power stations have continuous radiation monitoring and regular site inspections; new aircraft undergo extensive flight tests to prove their airworthiness. It’s no different for AI systems. New AI models go through a wide range of safety evaluations—for example, testing their capacity to assist in the creation of biological or chemical weapons. Such evaluations are built into our Responsible Scaling Policy, which guides our development of a model’s safeguards. As AIs become more capable, however, a new kind of risk might emerge: models with the ability to mislead their users, or subvert the systems we put in place to oversee them. A new paper by the Anthropic Alignment Science team describes a novel set of evaluations that test a model’s capacity for sabotage. We looked at four different types: We developed these evaluations as part of preparations for a future where much more capable AI models could develop these
Alignment Sabotage evaluations for frontier models Oct 18, 2024 Read the paper Any industry where there are potential harms needs evaluations. Nuclear power stations have continuous radiation monitoring and regular site inspections; new aircraft undergo extensive flight tests to prove their airworthiness. It’s no different for AI systems. New AI models go through a wide range of safety evaluations—for example, testing their capacity to assist in the creation of biological or chemical weapons. Such evaluations are built into our Responsible Scaling Policy , which guides our development of a mod
related reading
- 2410.21514arxiv.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Pre-deployment auditing can catch an overt saboteuralignment.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Toward A Public Science of Model Behavior | Transluce AItransluce.org
- 2312.06942arxiv.org
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Chapter 3: LLM Evaluations - ARENAlearn.arena.education
- Automated Researchers Can Subtly Sandbagalignment.anthropic.com