Simple probes can catch sleeper agents \ Anthropic
This “Alignment Note” presents some early-stage research from the Anthropic Alignment Science team following up on our recent “Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training” paper. It should be treated as a work-in-progress update, and is intended for a more technical audience than our typical blog post. This research makes use of some simple interpretability techniques, and we expect to share more results from collaborations between our Alignment and Interpretability teams soon. In this post we present “defection probes”: linear classifiers that use residual stream activations to predict when a sleeper agent trojan model will choose to “defect” and behave in accordance with a dangerous hidden goal. Using the models we trained in “Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training”, we show that linear detectors with AUROC scores above 99% can be created using generic contrast pairs that don’t depend on any information about the
Alignment Interpretability Simple probes can catch sleeper agents Apr 23, 2024 This “Alignment Note” presents some early-stage research from the Anthropic Alignment Science team following up on our recent “ Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training ” paper. It should be treated as a work-in-progress update, and is intended for a more technical audience than our typical blog post. This research makes use of some simple interpretability techniques, and we expect to share more results from collaborations between our Alignment and Interpretability teams soon. Sum
Explore this link on the map →related reading
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Detecting Strategic Deception Using Linear Probes — LessWronglesswrong.com
- secret-loyalties-whitepaper.pdfformationresearch.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Claude 4 System Cardwww-cdn.anthropic.com
- Frontier Risk Report (February to March 2026) - METRmetr.org