Neural Chameleons: LLMs Can Learn to Evade Activation Monitors
Language models can learn to evade activation monitors with 0-shot transfer to unseen deception and harmfulness detectors, while retaining capabilities.
Neural Chameleons Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors Authors: Max McGuinness (MATS), Alex Serrano (MATS, UC Berkeley), Luke Bailey (Stanford University), Scott Emmons (MATS) Summary We show that language models can learn a general skill to evade activation monitors, with 0-shot transfer to unseen deception and harmfulness monitors from the literature, with minimal effect on model capabilities. We call these "Neural Chameleons". Method We used a "stress-testing" approach: fine-tuning models to evade simple probes for benign concepts when given a tri
Explore this link on the map →related reading
- [2512.11949] Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitorsarxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Natural Language Autoencoders \ Anthropicanthropic.com
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- [2506.10805] Detecting High-Stakes Interactions with Activation Probesarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org