flâneur — a map of the web's best reading

Automated Researchers Can Subtly Sandbag

alignment.anthropic.com · 15,113 words · saved by 1 readers

When prompted, current models can sandbag ML experiments and research decisions without being detected by zero-shot prompted monitors. Claude 3.5 Sonnet (new) can only sandbag effectively when seeing a one-shot example, while Claude 3.7 Sonnet can do this without an example (zero-shot). We are not yet worried about sabotage in today's models, as the sandbagging we observe today would be noticed by humans.

Automated Researchers Can Subtly Sandbag Alignment Science Blog Automated Researchers Can Subtly Sandbag tl;dr When prompted, current models can sandbag ML experiments and research decisions without being detected by zero-shot prompted monitors. Claude 3.5 Sonnet (new) can only sandbag effectively when seeing a one-shot example, while Claude 3.7 Sonnet can do this without an example (zero-shot). We are not yet worried about sabotage in today's models, as the sandbagging we observe today would be noticed by humans. AI models might soon assist with large parts of AI and AI safety research, compl

Explore this link on the map →

saved by

related reading