You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWrong
lesswrong.com · 3,480 words · saved by 1 readers
Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning — preprint, 2026. [Paper] [Code] …
x You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWrong Interpretability (ML & AI) AI Frontpage 66 You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them by RobinHa 10th Jun 2026 Linkpost for robinhaselhorst.com 11 min read 5 66 Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning — preprint, 2026. [ Paper ] [ Code ] TLDR. Given a model with some unknown, abnormal behavior (backdoors, censorship, reward hacking, ...), construct an aligned reference by training a clean model to match the suspect's residual-stream activations on a beni
saved by
related reading
- [2510.04340] Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-timearxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- [2512.11949] Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitorsarxiv.org
- Discovering Backdoor Triggers — LessWronglesswrong.com
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2607.14111] Introspection Fine-Tuning (IFT): Training Small LLMs to Introspectarxiv.org
- Inside a Neural Chameleonjacksonmowattgok.com
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com