✳flâneur — a map of the web's best reading
You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWrong
lesswrong.com · 3,480 words · saved by 1 readers
Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning — preprint, 2026. [Paper] [Code] …
x You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWrong Interpretability (ML & AI) AI Frontpage 66 You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them by RobinHa 10th Jun 2026 Linkpost for robinhaselhorst.com 11 min read 5 66 Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning — preprint, 2026. [ Paper ] [ Code ] TLDR. Given a model with some unknown, abnormal behavior (backdoors, censorship, reward hacking, ...), construct an aligned reference by training a clean model to match the suspect's residual-stream activations on a beni
Explore this link on the map →saved by
related reading
- [2510.04340] Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-timearxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Discovering Backdoor Triggers — LessWronglesswrong.com
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- secret-loyalties-whitepaper.pdfformationresearch.com
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- Simple probes can catch sleeper agents \ Anthropicanthropic.com
- Backdoors have universal representations across large language models — LessWronglesswrong.com