Prediction, Explanation, or Over-interpretation?
Understanding and predicting the behavior of large language models remains a central challenge in interpretability and AI safety. Recent work has proposed natural language verbalization as a scalable approach for accessing model internals, including introspective prediction, activation-conditioned question answering, and reconstruction-based explanation (Binder et al., 2024, Li et al., 2025, Karvonen et al., 2025, Fraser-Taliente, Kantamneni, Ong et al., 2026). These approaches suggest that language models may be able to verbalize information about their own behaviors, latent states, or future generations. Despite promising results, it remains unclear what these verbalizations actually represent and whether different methods are even verbalizing the same thing. Each method optimizes a different objective, receives a different input, and is trained on different data, so a strict head-to-head comparison is difficult. To evaluate these verbalization methods, we examine whether natural-lan
Prediction, Explanation, or Over-interpretation? Prediction, Explanation, or Over-interpretation? Xiaoyan Bai Understanding and predicting the behavior of large language models remains a central challenge in interpretability and AI safety. Recent work has proposed natural language verbalization as a scalable approach for accessing model internals, including introspective prediction, activation-conditioned question answering, and reconstruction-based explanation ( Binder et al., 2024 , Li et al., 2025 , Karvonen et al., 2025 , Fraser-Taliente, Kantamneni, Ong et al., 2026 ). These approaches su
Explore this link on the map →saved by
related reading
- Can activation verbalizers surface an internal chain of thought? — LessWronglesswrong.com
- Transformer Circuits Threadtransformer-circuits.pub
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Language Models, World Models, and Human Model-Buildinglingo.csail.mit.edu
- Natural Language Autoencoders \ Anthropicanthropic.com
- [2509.13316] Do Activation Verbalization Methods Convey Privileged Information?arxiv.org
- [2509.13316] Do Activation Verbalization Methods Convey Privileged Information?arxiv.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- How do we know if activation verbalizers are telling us anything about activations? — Millicent Limillicentli.github.io