Research
Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation
arXiv:2608.19611v1 Announce Type: cross Abstract: LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produce for a giv
arXiv:2608.19611v1 Announce Type: cross Abstract: LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produce for a given question, i.e., its uncertainty. Resampling-based analyses characterize this distribution, revealing which steps of a rollout determine how the model arrives at its answer. However, a major limitation of these approaches is that resampling text sequences at every token or sentence in a reasoning chain is very costly. Our work strives to make resampling analysis more computationally efficient, while also shedding light on an important scientific question: what is the right statistical model for explaining uncertainty dynamics in text generation? We show that when resampling many reasoning chains, uncertainty dynamics converge to stable patterns, and noise is largely an artifact of sampling rather than an LLM's sensitivity to each individual token or reasoning step. We develop a statistical model for smoothing noisy low-sample rollout data to better approximate high-sample data, allowing us to significantly cut sampling costs.
Related
- Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
- TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
- EAGer: Entropy-Aware GEneRation for Adaptive Inference-Time Scaling
- Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
Source: arXiv cs.AI | 2026-08-21