Automated Researchers Can Subtly Sandbag
When prompted, current models can sandbag ML experiments and research decisions without being detected by zero-shot prompted monitors. Claude 3.5 Sonnet (new) can only sandbag effectively when seeing a one-shot example, while Claude 3.7 Sonnet can do this without an example (zero-shot). We are not yet worried about sabotage in today's models, as the sandbagging we observe today would be noticed by humans.
Automated Researchers Can Subtly Sandbag Alignment Science Blog Automated Researchers Can Subtly Sandbag tl;dr When prompted, current models can sandbag ML experiments and research decisions without being detected by zero-shot prompted monitors. Claude 3.5 Sonnet (new) can only sandbag effectively when seeing a one-shot example, while Claude 3.7 Sonnet can do this without an example (zero-shot). We are not yet worried about sabotage in today's models, as the sandbagging we observe today would be noticed by humans. AI models might soon assist with large parts of AI and AI safety research, compl
saved by
related reading
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Sandbagging with misaligned action - Chain-of-Thought Transcript - Anti-Schemingantischeming.ai
- Pre-deployment auditing can catch an overt saboteuralignment.anthropic.com
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Teaching Claude Whyalignment.anthropic.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Auditing Games for Sandbagging [paper] — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Claude 4 System Cardwww-cdn.anthropic.com