[2509.13316] Do Natural Language Descriptions of Model Activations Convey Privileged Information?
Abstract:Recent interpretability methods have proposed to translate LLM internal representations into natural language descriptions using a second verbalizer LLM. This is intended to illuminate how the target model represents and operates on inputs. But do such activation verbalization approaches actually provide privileged knowledge about the internal workings of the target model, or do they merely convey information about its inputs? We critically evaluate popular verbalization methods across datasets used in prior work and find that they can succeed at benchmarks without any access to target model internals, suggesting that these datasets may not be ideal for evaluating verbalization methods. We then run controlled experiments which reveal that verbalizations often reflect the parametric knowledge of the verbalizer LLM which generated them, rather than the knowledge of the target LLM whose activations are decoded. Taken together, our results indicate a need for targeted benchmarks and experimental controls to rigorously assess whether verbalization methods provide meaningful insights into the operations of LLMs.
[2509.13316] Do Activation Verbalization Methods Convey Privileged Information? Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Computation and Language arXiv:2509.13316 (cs) [Submitted on 16 Sep 2025 ( v1 ), last revised 13 May 2026 (this version, v4)] Title: Do Activation Verbalization Methods Convey Privileged Information? Authors: Millicent Li , Alberto Mario Ceballos Arroyo , Giordano Rogers , Naomi Saphra , Byron C. Wallace View a PDF of the paper titled Do Activation Verbali
Explore this link on the map →related reading
- [2509.13316] Do Activation Verbalization Methods Convey Privileged Information?arxiv.org
- How do we know if activation verbalizers are telling us anything about activations? — Millicent Limillicentli.github.io
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Can activation verbalizers surface an internal chain of thought? — LessWronglesswrong.com
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWronglesswrong.com
- [2512.15674] Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersarxiv.org
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersalignment.anthropic.com