Automated Researchers Can Subtly Sandbag
When prompted, current models can sandbag ML experiments and research decisions without being detected by zero-shot prompted monitors. Claude 3.5 Sonnet (new) can only sandbag effectively when seeing a one-shot example, while Claude 3.7 Sonnet can do this without an example (zero-shot). We are not yet worried about sabotage in today's models, as the sandbagging we observe today would be noticed by humans.
Automated Researchers Can Subtly Sandbag Alignment Science Blog Automated Researchers Can Subtly Sandbag tl;dr When prompted, current models can sandbag ML experiments and research decisions without being detected by zero-shot prompted monitors. Claude 3.5 Sonnet (new) can only sandbag effectively when seeing a one-shot example, while Claude 3.7 Sonnet can do this without an example (zero-shot). We are not yet worried about sabotage in today's models, as the sandbagging we observe today would be noticed by humans. AI models might soon assist with large parts of AI and AI safety research, compl
Explore this link on the map →saved by
related reading
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Sandbagging with misaligned action - Chain-of-Thought Transcript - Anti-Schemingantischeming.ai
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Auditing Games for Sandbagging [paper] — LessWronglesswrong.com
- Claude 4 System Cardwww-cdn.anthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWronglesswrong.com
- Thoughts on Claude Fable's silent safeguards — LessWronglesswrong.com
- Sabotage evaluations for frontier models \ Anthropicanthropic.com