flâneur — a map of the web's best reading

Predictive Concept Decoders | Transluce AI

transluce.org · 1,497 words · saved by 1 readers

We train interpretability assistants to look inside AI models and accurately answer questions about their behavior. These assistants, which we call Predictive Concept Decoders (PCDs), learn to translate a model's internal states into short, human-readable concept lists, then use those concepts to answer questions about the model's behavior. PCDs become both more accurate and more interpretable as we scale up training, and they can surface information that models hide or fail to report, including jailbreaks, secret hints, and implanted concepts.

Predictive Concept Decoders Training scalable end-to-end interpretability assistants Vincent Huang * , Dami Choi , Daniel D. Johnson , Sarah Schwettmann , Jacob Steinhardt * Correspondence to: vincent@transluce.org Transluce | Published: December 18, 2025 We train interpretability assistants to look inside AI models and accurately answer questions about their behavior. These assistants, which we call Predictive Concept Decoders (PCDs), learn to translate a model's internal states into short, human-readable concept lists, then use those concepts to answer questions about the model's behavior. P

Explore this link on the map →

related reading