goodfire.ai/blog/under-the-hood-of-a-reasoning-model
We have trained the first ever sparse autoencoders (SAEs) on the 671B parameter DeepSeek R1 model and open-sourced the SAEs. Dron Hazra † Max Loeffler † Murat Cubuktepe Levon Avagyan ‡ Liv Gorton Mark Bissell Owen Lewis Thomas McGrath Daniel Balsam * Apr. 15, 2025 Today, we’re excited to share some early work on mechanistic interpretability for DeepSeek’s reasoning model R1—including two state-of-the-art, open-sourced sparse autoencoders (SAEs). These are the first public interpreter models trained on a true reasoning model, and on any model of this scale. Our early experiments indicate that R1 is qualitatively different from non-reasoning language models, and required some novel insights to steer it. While we plan to do a lot more work to understand the internal workings of reasoning models such as R1, we’re excited to share these initial contributions with the research community. We have trained both a general reasoning and math specific SAE. Because R1 is a very large model and ther
Under the Hood of a Reasoning Model Research Under the Hood of a Reasoning Model We have trained the first ever sparse autoencoders (SAEs) on the 671B parameter DeepSeek R1 model and open-sourced the SAEs. Authors Dron Hazra † Max Loeffler † Murat Cubuktepe Levon Avagyan ‡ Liv Gorton Mark Bissell Owen Lewis Thomas McGrath Daniel Balsam * Published Apr. 15, 2025 † Core contributor ‡ Independent researcher, work done while visiting Goodfire * Correspondence to dan@goodfire.ai Today, we’re excited to share some early work on mechanistic interpretability for DeepSeek’s reasoning model R1—including
related reading
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- DeepSeek-R1arxiv.org
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Transformer Circuits Threadtransformer-circuits.pub
- A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Modelsarxiv.org
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- An Intuitive Explanation of Sparse Autoencoders for Mechanistic Interpretability of LLMs — LessWronglesswrong.com
- Base Models Know How to Reason, Thinking Models Learn Whenarxiv.org
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible.github.com
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io