Mapping global dynamics of benchmark creation and saturation in artificial intelligence | Nature Communications
Recent studies raised concerns over the state of AI benchmarking, reporting issues such as benchmark overfitting, benchmark saturation and increasing centralization of benchmark dataset creation. To facilitate monitoring of the health of the AI benchmarking ecosystem, the authors introduce methodologies for creating condensed maps of the global dynamics of benchmark.
Download PDF Subjects Computer science Scientific data Abstract Benchmarks are crucial to measuring and steering progress in artificial intelligence (AI). However, recent studies raised concerns over the state of AI benchmarking, reporting issues such as benchmark overfitting, benchmark saturation and increasing centralization of benchmark dataset creation. To facilitate monitoring of the health of the AI benchmarking ecosystem, we introduce methodologies for creating condensed maps of the global dynamics of benchmark creation and saturation. We curate data for 3765 benchmarks covering the ent
Explore this link on the map →related reading
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- AI as Normal Technology | Knight First Amendment Instituteknightcolumbia.org
- AI’s capabilities may be exaggerated by flawed tests, study saysnbcnews.com
- Import AIjack-clark.net
- AI Benchmarking Is Brokenwhoisnnamdi.com
- [1911.01547] On the Measure of Intelligencearxiv.org
- Preface - The Emerging Science of Machine Learning Benchmarksmlbenchmarks.org
- [2606.05405] Agents' Last Examarxiv.org
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- Devising ML Metrics | CAISsafe.ai
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- Research Areas in Benchmark Design and Evaluation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org