Steerling-8B: The First Inherently Interpretable Language Model
We release Steerling-8B, an 8B-parameter causal diffusion language model that is interpretable by construction — its predictions are routed through concepts you can measure, audit, and control.
Steerling-8B: The First Inherently Interpretable Language Model Steerling-8B: The First Inherently Interpretable Language Model Author: Guide Labs Team Published: February 23, 2026 We are releasing Steerling-8B, the first interpretable model that can trace any token it generates to its input context, concepts a human can understand, and its training data. Trained on 1.35 trillion tokens, the model achieves downstream performance within range of models trained on 2–7× more data. Steerling-8B unlocks several capabilities which include suppressing or amplifying specific concepts at inference time
Explore this link on the map →related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Transformer Circuits Threadtransformer-circuits.pub
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Should Developers Care about Interpretability? • Thariq Shihiparthariq.io
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- Tracing the Thoughts of a Large Language Model — LessWronglesswrong.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Tracing Model Outputs to the Training Data \ Anthropicanthropic.com
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- Mapping the mind of a large language model \ Anthropicanthropic.com