Prediction, Explanation, or Over-interpretation?
Understanding and predicting the behavior of large language models remains a central challenge in interpretability and AI safety. Recent work has proposed natural language verbalization as a scalable approach for accessing model internals, including introspective prediction, activation-conditioned question answering, and reconstruction-based explanation (Binder et al., 2024, Li et al., 2025, Karvonen et al., 2025, Fraser-Taliente, Kantamneni, Ong et al., 2026). These approaches suggest that language models may be able to verbalize information about their own behaviors, latent states, or future generations. Despite promising results, it remains unclear what these verbalizations actually represent and whether different methods are even verbalizing the same thing. Each method optimizes a different objective, receives a different input, and is trained on different data, so a strict head-to-head comparison is difficult. To evaluate these verbalization methods, we examine whether natural-lan
Prediction, Explanation, or Over-interpretation? Prediction, Explanation, or Over-interpretation? Xiaoyan Bai Understanding and predicting the behavior of large language models remains a central challenge in interpretability and AI safety. Recent work has proposed natural language verbalization as a scalable approach for accessing model internals, including introspective prediction, activation-conditioned question answering, and reconstruction-based explanation ( Binder et al., 2024 , Li et al., 2025 , Karvonen et al., 2025 , Fraser-Taliente, Kantamneni, Ong et al., 2026 ). These approaches su
saved by
related reading
- Can activation verbalizers surface an internal chain of thought? — LessWronglesswrong.com
- Transformer Circuits Threadtransformer-circuits.pub
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- Natural Language Autoencoders \ Anthropicanthropic.com
- [2511.08579] Training Language Models to Explain Their Own Computationsarxiv.org
- Emergent introspective awareness in large language models \ Anthropicanthropic.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2509.13316] Do Activation Verbalization Methods Convey Privileged Information?arxiv.org
- Training Language Models to Explain Their Own Computationsarxiv.org
- [2509.13316] Do Activation Verbalization Methods Convey Privileged Information?arxiv.org