Cycles of Thought: Measuring LLM Confidence through Stable Explanations
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. \ul In many critical machine learning (ML) applications it is essential for a model to indicate when it is uncertain about a prediction. While large language models (LLMs) can reach and even surpass human-level accuracy on a variety of benchmarks, their overconfidence in incorrect responses is still a well-documented failure mode. Traditional methods for ML uncertainty quantification can be difficult to directly adapt to LLMs due to the computational cost of implementation and closed-source nature of many models. A variety of black-box methods have recently been proposed, but these often rely on heuristics such as self-verbalized confidence. We instead propose a framework for measuring an LLM’s uncertainty with respect to the distribution of generated ex
\useunder \ul Cycles of Thought: Measuring LLM Confidence through Stable Explanations Evan Becker Department of Computer Science UCLA evbecker@ucla.edu &Stefano Soatto Department of Computer Science UCLA soatto@cs.ucla.edu (October 16, 2025) Abstract In many critical machine learning (ML) applications it is essential for a model to indicate when it is uncertain about a prediction. While large language models (LLMs) can reach and even surpass human-level accuracy on a variety of benchmarks, their overconfidence in incorrect responses is still a well-documented failure mode. Traditional methods
Explore this link on the map →related reading
- [2405.20974] SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationalesarxiv.org
- [2405.20974] SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationalesarxiv.org
- [2305.04388] Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Promptingarxiv.org
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Detecting when LLMs are Uncertain • Thariq Shihiparthariq.io
- [2602.02639] A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behaviorarxiv.org
- Applying Statistics to LLM Evaluationscameronrwolfe.substack.com
- How confessions can keep language models honest | OpenAIopenai.com
- Taking LLMs Seriously (As Language Models) — LessWronglesswrong.com