flâneur — a map of the web's best reading

How will we do SFT on models with opaque reasoning? — AI Alignment Forum

alignmentforum.org · 2,090 words · saved by 1 readers

Current LLMs externalize lots of their reasoning in human interpretable language. This reasoning is sometimes unfaithful, sometimes strange and concerning, and LLMs can do somewhat impressive reasoning without using CoT, but my overall impression is that CoT currently is a reasonably complete and accurate representation of LLM reasoning. However, reasoning in interpretable language might turn out to be uncompetitive—if so, it seems probable that opaque reasoning will be adopted in frontier AI labs. If future AI models have opaque reasoning, this will probably change what training we can apply to these AIs. For example, currently we train models to reason in a good way about math problems, or to reason in a desired way about the spec that we hope they’ll follow. It’s not obvious that we’ll be able to do training that affects model reasoning like this if models have opaque reasoning though, because we can’t just write the reasoning ourselves and do SFT on the reasoning trace. Prior work

x How will we do SFT on models with opaque reasoning? — AI Alignment Forum Considerations in diffuse control AI Control AI Frontpage 18 How will we do SFT on models with opaque reasoning? by Alek Westover , Vivek Hebbar , egan 21st Feb 2026 8 min read 17 18 Current LLMs externalize lots of their reasoning in human interpretable language. This reasoning is sometimes unfaithful , sometimes strange and concerning , and LLMs can do somewhat impressive reasoning without using CoT , but my overall impression is that CoT currently is a reasonably complete and accurate representation of LLM reasoning.

Explore this link on the map →

related reading