How do we know if activation verbalizers are telling us anything about activations? — Millicent Li
TL;DR: Evaluations for activation verbalizers should establish simple baselines to inform us whether these verbalizers can actually yield information about the target LM, beyond what is accessible with black-box baselines. We find that evaluations for recent activation verbalizers, like Anthropic’s Natural Language Autoencoder (NLA), do not sufficiently meet this bar. Recently, there has been excitement around activation verbalization[1]: decoding the activations (or intermediate hidden states) of a target language model (LM) into natural language with another verbalizer LM (or similar model)[3, 4, 5, 6, 7, 8, 9, 10, 11]. The great hope of activation verbalization is to access a LM's "inner thoughts": Thoughts that LMs might never articulate outright. For instance, these thoughts should be knowledge available to the target LM but not directly accessible to the verbalizer without access to the target LM's internals. Activation verbalizers (also known under different names, such as Activ
How do we know if activation verbalizers are telling us anything about activations? — Millicent Li ← back to blog How do we know if activation verbalizers are telling us anything about activations? July 1st, 2026 TL;DR: Evaluations for activation verbalizers should establish simple baselines to inform us whether these verbalizers can actually yield information about the target LM, beyond what is accessible with black-box baselines. We find that evaluations for recent activation verbalizers, like Anthropic’s Natural Language Autoencoder (NLA), do not sufficiently meet this bar. Figure 1: I
Explore this link on the map →related reading
- Can activation verbalizers surface an internal chain of thought? — LessWronglesswrong.com
- [2509.13316] Do Activation Verbalization Methods Convey Privileged Information?arxiv.org
- [2509.13316] Do Activation Verbalization Methods Convey Privileged Information?arxiv.org
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Natural Language Autoencoders \ Anthropicanthropic.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWronglesswrong.com
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersalignment.anthropic.com
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers — LessWronglesswrong.com
- Transformer Circuits Threadtransformer-circuits.pub
- [2512.15674] Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersarxiv.org