goodfire.ai/blog/under-the-hood-of-a-reasoning-model
We have trained the first ever sparse autoencoders (SAEs) on the 671B parameter DeepSeek R1 model and open-sourced the SAEs. Dron Hazra † Max Loeffler † Murat Cubuktepe Levon Avagyan ‡ Liv Gorton Mark Bissell Owen Lewis Thomas McGrath Daniel Balsam * Apr. 15, 2025 Today, we’re excited to share some early work on mechanistic interpretability for DeepSeek’s reasoning model R1—including two state-of-the-art, open-sourced sparse autoencoders (SAEs). These are the first public interpreter models trained on a true reasoning model, and on any model of this scale. Our early experiments indicate that R1 is qualitatively different from non-reasoning language models, and required some novel insights to steer it. While we plan to do a lot more work to understand the internal workings of reasoning models such as R1, we’re excited to share these initial contributions with the research community. We have trained both a general reasoning and math specific SAE. Because R1 is a very large model and ther
Under the Hood of a Reasoning Model Research Under the Hood of a Reasoning Model We have trained the first ever sparse autoencoders (SAEs) on the 671B parameter DeepSeek R1 model and open-sourced the SAEs. Authors Dron Hazra † Max Loeffler † Murat Cubuktepe Levon Avagyan ‡ Liv Gorton Mark Bissell Owen Lewis Thomas McGrath Daniel Balsam * Published Apr. 15, 2025 † Core contributor ‡ Independent researcher, work done while visiting Goodfire * Correspondence to dan@goodfire.ai Today, we’re excited to share some early work on mechanistic interpretability for DeepSeek’s reasoning model R1—including
Explore this link on the map →related reading
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- DeepSeek-R1arxiv.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Transformer Circuits Threadtransformer-circuits.pub
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- An Intuitive Explanation of Sparse Autoencoders for Mechanistic Interpretability of LLMs — LessWronglesswrong.com
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- Neuronpedianeuronpedia.org
- pdfopenreview.net