flâneur — a map of the web's best reading

You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWrong

lesswrong.com · 3,480 words · saved by 1 readers

Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning — preprint, 2026. [Paper] [Code] …

x You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWrong Interpretability (ML & AI) AI Frontpage 66 You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them by RobinHa 10th Jun 2026 Linkpost for robinhaselhorst.com 11 min read 5 66 Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning — preprint, 2026. [ Paper ] [ Code ] TLDR. Given a model with some unknown, abnormal behavior (backdoors, censorship, reward hacking, ...), construct an aligned reference by training a clean model to match the suspect's residual-stream activations on a beni

Explore this link on the map →

saved by

related reading