Measuring what Matters: Construct Validity in Large Language Model Benchmarks | alphaXiv
alphaxiv.org · 1,209 words · saved by 1 readers
Researchers systematically reviewed 445 Large Language Model benchmarks to assess construct validity, uncovering widespread issues in how abstract capabili
Understanding Construct Validity in Large Language Model Evaluation Large Language Models (LLMs) are increasingly evaluated using benchmarks that claim to measure complex capabilities like reasoning, safety, and intelligence. However, a fundamental question remains largely unexamined: do these benchmarks actually measure what they claim to measure? This paper addresses this critical gap by conducting the first large-scale systematic review of construct validity in LLM benchmarks, analyzing 445 peer-reviewed articles to understand current evaluation practices and their limitations.…
saved by
related reading
- Measuring what Matters: Construct Validity in Large Language Model Benchmarksalphaxiv.org
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Successful language model evals - Jason Weijasonwei.net
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- AI’s capabilities may be exaggerated by flawed tests, study saysnbcnews.com
- [2412.00543] Evaluating the Consistency of LLM Evaluatorsarxiv.org
- The bitter lesson of LLM evalsparsed.com
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- LLM evaluation: a beginner's guideevidentlyai.com
- [2506.13023] A Practical Guide for Evaluating LLMs and LLM-Reliant Systemsarxiv.org
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- 2410.05229arxiv.org