[2608.19611] Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation
Abstract:LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produce for a given question, i.e., its uncertainty. Resampling-based analyses characterize this distribution, revealing which steps of a rollout determine how the model arrives at its answer. However, a major limitation of these approaches is that resampling text sequences at every token or sentence in a reasoning chain is very costly. Our work strives to make resampling analysis more computationally efficient, while also shedding light on an important scientific question: what is the right statistical model for explaining uncertainty dynamics in text generation? We show that when resampling many reasoning chains, uncertainty dynamics converge to stable patterns, and noise is largely an artifact of sampling rather than an LLM's sensitivity to each individual token or reasoning step. We develop a statistical model for smoothing noisy low-sample rollout data to better approximate high-sample data, allowing us to significantly cut sampling costs.
View PDF HTML (experimental) Abstract:LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produce for a given question, i.e., its uncertainty. Resampling-based analyses characterize this distribution, revealing which steps of a rollout determine how the model arrives at its answer. However, a major limitation of these approaches is that resampling text sequences at every token or sentence in a reasoning chain is very costly. Our work strives to make resampling analysis more computationally efficient, while also…
saved by
related reading
- [2603.07267] How to Steal Reasoning Without Reasoning Tracesarxiv.org
- As Rocks May Think | Eric Jangevjang.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Thought Branches: Interpreting LLM Reasoning Requires Resamplingarxiv.org
- Large Language Models Must Be Taught to Know What They Don't Knowarxiv.org
- [2510.27484] Thought Branches: Interpreting LLM Reasoning Requires Resamplingarxiv.org
- GenAI Handbookgenai-handbook.github.io
- Detecting when LLMs are Uncertain • Thariq Shihiparthariq.io
- Towards a Typology of Strange LLM Chains-of-Thought1a3orn.com
- 2409.02908arxiv.org
- [2203.14465] STaR: Bootstrapping Reasoning With Reasoningarxiv.org
- Tenobrus (@tenobrus) on Xx.com