How will we do SFT on models with opaque reasoning? — AI Alignment Forum
Current LLMs externalize lots of their reasoning in human interpretable language. This reasoning is sometimes unfaithful, sometimes strange and concerning, and LLMs can do somewhat impressive reasoning without using CoT, but my overall impression is that CoT currently is a reasonably complete and accurate representation of LLM reasoning. However, reasoning in interpretable language might turn out to be uncompetitive—if so, it seems probable that opaque reasoning will be adopted in frontier AI labs. If future AI models have opaque reasoning, this will probably change what training we can apply to these AIs. For example, currently we train models to reason in a good way about math problems, or to reason in a desired way about the spec that we hope they’ll follow. It’s not obvious that we’ll be able to do training that affects model reasoning like this if models have opaque reasoning though, because we can’t just write the reasoning ourselves and do SFT on the reasoning trace. Prior work
x How will we do SFT on models with opaque reasoning? — AI Alignment Forum Considerations in diffuse control AI Control AI Frontpage 18 How will we do SFT on models with opaque reasoning? by Alek Westover , Vivek Hebbar , egan 21st Feb 2026 8 min read 17 18 Current LLMs externalize lots of their reasoning in human interpretable language. This reasoning is sometimes unfaithful , sometimes strange and concerning , and LLMs can do somewhat impressive reasoning without using CoT , but my overall impression is that CoT currently is a reasonably complete and accurate representation of LLM reasoning.
Explore this link on the map →related reading
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- DeepSeek-R1arxiv.org
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Worries about latent reasoning in LLMs — EA Forumforum.effectivealtruism.org
- o1 and Reasoning | AndoLogsblog.ando.ai
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- Thought Branches: Interpreting LLM Reasoning Requires Resamplingarxiv.org
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Announcing ReasoningLens — Visualizing and Diagnosing LLM Reasoning at a Glancehuggingface.co
- [2603.05706] Reasoning Models Struggle to Control their Chains of Thoughtarxiv.org