[2411.00640] Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Abstract:Evaluations are critical for understanding the capabilities of large language models (LLMs). Fundamentally, evaluations are experiments; but the literature on evaluations has largely ignored the literature from other sciences on experiment analysis and planning. This article shows researchers with some training in statistics how to think about and analyze data from language model evaluations. Conceptualizing evaluation questions as having been drawn from an unseen super-population, we present formulas for analyzing evaluation data, measuring differences between two models, and planning an evaluation experiment. We make a number of specific recommendations for running language model evaluations and reporting experiment results in a way that minimizes statistical noise and maximizes informativeness.
Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations Evan Miller Anthropic evanmiller@anthropic.com arXiv:2411.00640v1 [stat.AP] 1 Nov 2024 November 4, 2024…
saved by
related reading
- Applying Statistics to LLM Evaluationscameronrwolfe.substack.com
- A statistical approach to model evaluations \ Anthropicanthropic.com
- [2412.00543] Evaluating the Consistency of LLM Evaluatorsarxiv.org
- The bitter lesson of LLM evalsparsed.com
- GitHub - open-compass/opencompass: OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.github.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- Successful language model evals - Jason Weijasonwei.net
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Large Language Model: world models or surface statistics?thegradient.pub