[2305.14251] FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
Abstract:Evaluating the factuality of long-form text generated by large language models (LMs) is non-trivial because (1) generations often contain a mixture of supported and unsupported pieces of information, making binary judgments of quality inadequate, and (2) human evaluation is time-consuming and costly. In this paper, we introduce FActScore (Factual precision in Atomicity Score), a new evaluation that breaks a generation into a series of atomic facts and computes the percentage of atomic facts supported by a reliable knowledge source. We conduct an extensive human evaluation to obtain FActScores of people biographies generated by several state-of-the-art commercial LMs -- InstructGPT, ChatGPT, and the retrieval-augmented PerplexityAI -- and report new analysis demonstrating the need for such a fine-grained score (e.g., ChatGPT only achieves 58%). Since human evaluation is costly, we also introduce an automated model that estimates FActScore, using retrieval and a strong language model, with less than a 2% error rate. Finally, we use this automated metric to evaluate 6,500 generations from a new set of 13 recent LMs that would have cost $26K if evaluated by humans, with various findings: GPT-4 and ChatGPT are more factual than public models, and Vicuna and Alpaca are some of the best public models.
[2305.14251] FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation --> Computer Science > Computation and Language arXiv:2305.14251 (cs) [Submitted on 23 May 2023 ( v1 ), last revised 11 Oct 2023 (this version, v2)] Title: FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation Authors: Sewon Min , Kalpesh Krishna , Xinxi Lyu , Mike Lewis , Wen-tau Yih , Pang Wei Koh , Mohit Iyyer , Luke Zettlemoyer , Hannaneh Hajishirzi View a PDF of the paper titled FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long
Explore this link on the map →saved by
related reading
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- gpt-4.pdfcdn.openai.com
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Believe It or Not: How Deeply do LLMs Believe Implanted Facts?alignment.anthropic.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- The bitter lesson of LLM evalsparsed.com
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- GitHub - salesforce/AuditNLG: AuditNLG: Auditing Generative AI Language Modeling for Trustworthiness · GitHubgithub.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2005.11401] Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasksarxiv.org
- 2025.acl-long.896.pdfaclanthology.org