[2406.04244] Benchmark Data Contamination of Large Language Models: A Survey
Abstract:The rapid development of Large Language Models (LLMs) like GPT-4, Claude-3, and Gemini has transformed the field of natural language processing. However, it has also resulted in a significant issue known as Benchmark Data Contamination (BDC). This occurs when language models inadvertently incorporate evaluation benchmark information from their training data, leading to inaccurate or unreliable performance during the evaluation phase of the process. This paper reviews the complex challenge of BDC in LLM evaluation and explores alternative assessment methods to mitigate the risks associated with traditional benchmarks. The paper also examines challenges and future directions in mitigating BDC risks, highlighting the complexity of the issue and the need for innovative solutions to ensure the reliability of LLM evaluation in real-world applications.
Abstract:The rapid development of Large Language Models (LLMs) like GPT-4, Claude-3, and Gemini has transformed the field of natural language processing. However, it has also resulted in a significant issue known as Benchmark Data Contamination (BDC). This occurs when language models inadvertently incorporate evaluation benchmark information from their training data, leading to inaccurate or unreliable performance during the evaluation phase of the process. This paper reviews the complex challenge of BDC in LLM evaluation and explores alternative assessment methods to mitigate the risks associ
Explore this link on the map →related reading
- gpt-4.pdfcdn.openai.com
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- The bitter lesson of LLM evalsparsed.com
- [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- [2506.13023] A Practical Guide for Evaluating LLMs and LLM-Reliant Systemsarxiv.org
- Questionable practices in machine learningarxiv.org
- Things we learned about LLMs in 2024simonwillison.net
- Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- Lapis Labslapis.rocks
- LLM evaluation: a beginner's guideevidentlyai.com