flâneur — a map of the web's best reading

Vestigial reasoning in RL — LessWrong

lesswrong.com · 4,576 words · saved by 1 readers

TL;DR: I claim that many reasoning patterns that appear in chains-of-thought are not actually used by the model to come to its answer, and can be more accurately thought of as historical artifacts of training. This can be true even for CoTs that are apparently "faithful" to the true reasons for the model's answer. Epistemic status: I'm pretty confident that the model described here is more accurate than my previous understanding. However, I wouldn't be very surprised if parts of this post are significantly wrong or misleading. Further experiments would be helpful for validating some of these hypotheses. Thanks to @Andy Arditi and @David Lindner for giving feedback on a draft of this post. Until recently, I assumed that RL training would cause reasoning models to make their chains-of-thought as efficient as possible, so that every token is directly useful to the model. However, I now believe that by default,[1] reasoning models' CoTs will often include many "useless" tokens that don't h

x Vestigial reasoning in RL — LessWrong AI Frontpage 69 Vestigial reasoning in RL by Caleb Biddulph 13th Apr 2025 11 min read 8 69 TL;DR: I claim that many reasoning patterns that appear in chains-of-thought are not actually used by the model to come to its answer, and can be more accurately thought of as historical artifacts of training. This can be true even for CoTs that are apparently "faithful" to the true reasons for the model's answer. Epistemic status: I'm pretty confident that the model described here is more accurate than my previous understanding. However, I wouldn't be very surpris

Explore this link on the map →

related reading