Base Models Know How to Reason, Thinking Models Learn When
What do thinking language models learn during training that their base models lack? We first present an unsupervised method that discovers a model’s reasoning behaviors by training small Sparse Autoencoders on sentence-level activations of reasoning traces, yielding interpretable reasoning taxonomies. Building on this, we introduce constructive model diffing, which aims to reconstruct the base-to-fine-tuned difference from interpretable components: reasoning mechanisms (category vectors that can induce a reasoning behavior in the base model) and reasoning heuristics (a classifier determining when a mechanism should fire). Across nine base/thinking pairs (four RL-trained, four SFT-distilled, one mixed), two independent findings agree: category vectors in the base model converge to far lower loss for taxonomies derived from purely RL-trained models, and hybrid models recover roughly 76% of the RL base-to-thinking gap but only 11% of the SFT gap. This indicates RL primarily teaches heuris
Constantin Venhoff Affiliation: University of Oxford, UK Affiliation: MATS Correspondence to: constantin@robots.ox.ac.uk Philip Torr Affiliation: University of Oxford, UK Arthur Conmy Neel Nanda Abstract What do thinking language models learn during training that their base models lack? We first present an unsupervised method that discovers a model’s reasoning behaviors by training small Sparse Autoencoders on sentence-level activations of reasoning traces, yielding interpretable reasoning taxonomies. Building on this, we introduce constructive model diffing, which aims to…
saved by
related reading
- Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Modelsarxiv.org
- DeepSeek-R1arxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- As Rocks May Think | Eric Jangevjang.com
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- Explore | alphaXivalphaxiv.org
- [2603.07267] How to Steal Reasoning Without Reasoning Tracesarxiv.org
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- Training Large Language Models to Reason in a Continuous Latent Spacearxiv.org
- [2504.13837] Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?arxiv.org
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexitymachinelearning.apple.com