[2406.08391] Large Language Models Must Be Taught to Know What They Don't Know
Abstract:When using large language models (LLMs) in high-stakes applications, we need to know when we can trust their predictions. Some works argue that prompting high-performance LLMs is sufficient to produce calibrated uncertainties, while others introduce sampling methods that can be prohibitively expensive. In this work, we first argue that prompting on its own is insufficient to achieve good calibration and then show that fine-tuning on a small dataset of correct and incorrect answers can create an uncertainty estimate with good generalization and small computational overhead. We show that a thousand graded examples are sufficient to outperform baseline methods and that training through the features of a model is necessary for good performance and tractable for large open-source models when using LoRA. We also investigate the mechanisms that enable reliable LLM uncertainty estimation, finding that many models can be used as general-purpose uncertainty estimators, applicable not just to their own uncertainties but also the uncertainty of other models. Lastly, we show that uncertainty estimates inform human use of LLMs in human-AI collaborative settings through a user study.
View PDF HTML (experimental) Abstract:When using large language models (LLMs) in high-stakes applications, we need to know when we can trust their predictions. Some works argue that prompting high-performance LLMs is sufficient to produce calibrated uncertainties, while others introduce sampling methods that can be prohibitively expensive. In this work, we first argue that prompting on its own is insufficient to achieve good calibration and then show that fine-tuning on a small dataset of correct and incorrect answers can create an uncertainty estimate with good generalization and small…
saved by
related reading
- [2302.00805] Conditioning Predictive Models: Risks and Strategiesarxiv.org
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- Productizing Large Language Modelsblog.replit.com
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Large Language Model: world models or surface statistics?thegradient.pub
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- Cycles of Thought: Measuring LLM Confidence through Stable Explanationsarxiv.org
- Things we learned about LLMs in 2024simonwillison.net
- 2023.emnlp-main.330.pdfaclanthology.org
- Stuff we figured out about AI in 2023simonwillison.net
- Foundational Challenges in Assuring Alignment and Safety of Large Language Modelsarxiv.org
- [2605.27288] It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertaintyarxiv.org