[2309.11495] Chain-of-Verification Reduces Hallucination in Large Language Models
Abstract:Generation of plausible yet incorrect factual information, termed hallucination, is an unsolved issue in large language models. We study the ability of language models to deliberate on the responses they give in order to correct their mistakes. We develop the Chain-of-Verification (CoVe) method whereby the model first (i) drafts an initial response; then (ii) plans verification questions to fact-check its draft; (iii) answers those questions independently so the answers are not biased by other responses; and (iv) generates its final verified response. In experiments, we show CoVe decreases hallucinations across a variety of tasks, from list-based questions from Wikidata, closed book MultiSpanQA and longform text generation.
View PDF HTML (experimental) Abstract:Generation of plausible yet incorrect factual information, termed hallucination, is an unsolved issue in large language models. We study the ability of language models to deliberate on the responses they give in order to correct their mistakes. We develop the Chain-of-Verification (CoVe) method whereby the model first (i) drafts an initial response; then (ii) plans verification questions to fact-check its draft; (iii) answers those questions independently so the answers are not biased by other responses; and (iv) generates its final verified response.…
saved by
related reading
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- (Im)possibility of Automated Hallucination Detection in Large Language Modelsarxiv.org
- Hallucination Mitigation using Agentic AI Natural Language-Based Frameworksarxiv.org
- Real-Time Detection of Hallucinated Entities in Long-Form Generationhallucination-probes.com
- HALVA: Hallucination Attenuated Language and Vision Assistantresearch.google
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Thinking to recall: How reasoning unlocks parametric knowledge in LLMsresearch.google
- MMHal Bencharxiv.org
- [2603.07267] How to Steal Reasoning Without Reasoning Tracesarxiv.org