Steerling-8B: The First Inherently Interpretable Language Model
We release Steerling-8B, an 8B-parameter causal diffusion language model that is interpretable by construction — its predictions are routed through concepts you can measure, audit, and control.
Steerling-8B: The First Inherently Interpretable Language Model Steerling-8B: The First Inherently Interpretable Language Model Author: Guide Labs Team Published: February 23, 2026 We are releasing Steerling-8B, the first interpretable model that can trace any token it generates to its input context, concepts a human can understand, and its training data. Trained on 1.35 trillion tokens, the model achieves downstream performance within range of models trained on 2–7× more data. Steerling-8B unlocks several capabilities which include suppressing or amplifying specific concepts at inference time
Explore this link on the map →related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Transformer Circuits Threadtransformer-circuits.pub
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- [2511.08579] Training Language Models to Explain Their Own Computationsarxiv.org
- Training Language Models to Explain Their Own Computationsarxiv.org
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- Training Language Models to Explain Their Own Computationsarxiv.org
- [2607.14111] Introspection Fine-Tuning (IFT): Training Small LLMs to Introspectarxiv.org
- Should Developers Care about Interpretability? • Thariq Shihiparthariq.io
- Topicslearnmechinterp.com
- What language do language models speak?tcz.hu
- Faithful, Interpretable Model Explanations via Causal Abstraction | SAIL Blogai.stanford.edu