Predictive Concept Decoders | Transluce AI
We train interpretability assistants to look inside AI models and accurately answer questions about their behavior. These assistants, which we call Predictive Concept Decoders (PCDs), learn to translate a model's internal states into short, human-readable concept lists, then use those concepts to answer questions about the model's behavior. PCDs become both more accurate and more interpretable as we scale up training, and they can surface information that models hide or fail to report, including jailbreaks, secret hints, and implanted concepts.
Predictive Concept Decoders Training scalable end-to-end interpretability assistants Vincent Huang * , Dami Choi , Daniel D. Johnson , Sarah Schwettmann , Jacob Steinhardt * Correspondence to: vincent@transluce.org Transluce | Published: December 18, 2025 We train interpretability assistants to look inside AI models and accurately answer questions about their behavior. These assistants, which we call Predictive Concept Decoders (PCDs), learn to translate a model's internal states into short, human-readable concept lists, then use those concepts to answer questions about the model's behavior. P
related reading
- [2512.15712] Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistantsarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Transformer Circuits Threadtransformer-circuits.pub
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Natural Language Autoencoders \ Anthropicanthropic.com
- Neuronpedianeuronpedia.org
- Goodfire AIgoodfire.ai
- Topicslearnmechinterp.com
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io