Neural Chameleons: LLMs Can Learn to Evade Activation Monitors
Language models can learn to evade activation monitors with 0-shot transfer to unseen deception and harmfulness detectors, while retaining capabilities.
Neural Chameleons Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors Authors: Max McGuinness (MATS), Alex Serrano (MATS, UC Berkeley), Luke Bailey (Stanford University), Scott Emmons (MATS) Summary We show that language models can learn a general skill to evade activation monitors, with 0-shot transfer to unseen deception and harmfulness monitors from the literature, with minimal effect on model capabilities. We call these "Neural Chameleons". Method We used a "stress-testing" approach: fine-tuning models to evade simple probes for benign concepts when given a tri
related reading
- [2512.11949] Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitorsarxiv.org
- Inside a Neural Chameleonjacksonmowattgok.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- [2608.21664] Measuring Activation Control in Large Language Modelsarxiv.org
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2607.14111] Introspection Fine-Tuning (IFT): Training Small LLMs to Introspectarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- 2212.03827arxiv.org