Inside a Neural Chameleon
jacksonmowattgok.com · 4,289 words · saved by 3 readers
Reverse-engineering how a language model hides from activation monitors
A language model that learned to hide# Neural Chameleons showed that models can learn a general skill for altering their neural activity to hide from completely new, unseen activation monitors.[1]Mainly tested using unseen linear probes but evasion also carried to some probe ensembles, MLPs, and attention monitors. How is this even possible? This seems like a bizarre meta-skill to be capable of. McGuinness et al. describes that: “When triggered, the model pushes its activations into a narrow, low-dimensional subspace, effectively steering them away from the regions a probe would flag as…
saved by
related reading
- [2512.11949] Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitorsarxiv.org
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Scaling Activation Oracles to Trillion-Parameter Modelstransluce.org
- Neural Chameleons: LLMs Can Learn to Evade Activation Monitorsneuralchameleons.com
- [2608.21664] Measuring Activation Control in Large Language Modelsarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Natural Language Autoencoders \ Anthropicanthropic.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- [2606.04071] Covert Influence Between Language Modelsarxiv.org