LatentQA: Teaching LLMs to Decode Activations Into Natural Language
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. Interpretability methods seek to understand language model representations, yet the outputs of most such methods—circuits, vectors, scalars—are not immediately human-interpretable. In response, we introduce LatentQA, the task of answering open-ended questions about model activations in natural language. Towards solving LatentQA, we propose Latent Interpretation Tuning (Lit),
LatentQA: Teaching LLMs to Decode Activations Into Natural Language Alexander Pan UC Berkeley &Lijie Chen UC Berkeley&Jacob Steinhardt UC Berkeley Correspondence to aypan.17@berkeley.edu . Project page: https://latentqa.github.io Abstract Interpretability methods seek to understand language model representations, yet the outputs of most such methods—circuits, vectors, scalars—are not immediately human-interpretable. In response, we introduce LatentQA , the task of answering open-ended questions about model activations in natural language. Towards solving LatentQA , we propose Latent Interpreta
Explore this link on the map →related reading
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Natural Language Autoencoders \ Anthropicanthropic.com
- [2512.15674] Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersarxiv.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWronglesswrong.com
- Transformer Circuits Threadtransformer-circuits.pub
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers — LessWronglesswrong.com
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainersalignment.anthropic.com
- Neuronpedianeuronpedia.org
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org