Cycles of Thought: Measuring LLM Confidence through Stable Explanations
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. \ul In many critical machine learning (ML) applications it is essential for a model to indicate when it is uncertain about a prediction. While large language models (LLMs) can reach and even surpass human-level accuracy on a variety of benchmarks, their overconfidence in incorrect responses is still a well-documented failure mode. Traditional methods for ML uncertainty quantification can be difficult to directly adapt to LLMs due to the computational cost of implementation and closed-source nature of many models. A variety of black-box methods have recently been proposed, but these often rely on heuristics such as self-verbalized confidence. We instead propose a framework for measuring an LLM’s uncertainty with respect to the distribution of generated ex
\useunder \ul Cycles of Thought: Measuring LLM Confidence through Stable Explanations Evan Becker Department of Computer Science UCLA evbecker@ucla.edu &Stefano Soatto Department of Computer Science UCLA soatto@cs.ucla.edu (October 16, 2025) Abstract In many critical machine learning (ML) applications it is essential for a model to indicate when it is uncertain about a prediction. While large language models (LLMs) can reach and even surpass human-level accuracy on a variety of benchmarks, their overconfidence in incorrect responses is still a well-documented failure mode. Traditional methods
related reading
- Large Language Models Must Be Taught to Know What They Don't Knowarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- 2023.emnlp-main.330.pdfaclanthology.org
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- [2511.08579] Training Language Models to Explain Their Own Computationsarxiv.org
- confessions_paper.pdfcdn.openai.com
- [2605.27288] It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertaintyarxiv.org
- [2405.20974] SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationalesarxiv.org
- Self-CTRL: Self-Consistency Training with Reinforcement Learningarxiv.org
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- Position: It's Time to Optimize for Self-Consistencytime-for-consistency.github.io