Predictive Concept Decoders | Transluce AI
We train interpretability assistants to look inside AI models and accurately answer questions about their behavior. These assistants, which we call Predictive Concept Decoders (PCDs), learn to translate a model's internal states into short, human-readable concept lists, then use those concepts to answer questions about the model's behavior. PCDs become both more accurate and more interpretable as we scale up training, and they can surface information that models hide or fail to report, including jailbreaks, secret hints, and implanted concepts.
Predictive Concept Decoders Training scalable end-to-end interpretability assistants Vincent Huang * , Dami Choi , Daniel D. Johnson , Sarah Schwettmann , Jacob Steinhardt * Correspondence to: vincent@transluce.org Transluce | Published: December 18, 2025 We train interpretability assistants to look inside AI models and accurately answer questions about their behavior. These assistants, which we call Predictive Concept Decoders (PCDs), learn to translate a model's internal states into short, human-readable concept lists, then use those concepts to answer questions about the model's behavior. P
Explore this link on the map →related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- [2512.15712] Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistantsarxiv.org
- Transformer Circuits Threadtransformer-circuits.pub
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Neuronpedianeuronpedia.org
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible. · GitHubgithub.com
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Scalable End-to-End Interpretability — LessWronglesswrong.com