Call For Evaluations & Datasets 2026
The NeurIPS Evaluations & Datasets (E&D) Track invites submissions that advance the science and practice of evaluation in AI/ML including the development and use of datasets and other resources. Formerly the Datasets & Benchmarks Track, the E&D Track reflects a shift and broadening in scope: evaluation becomes an object of scientific study in its own right. Scientific debates increasingly hinge on evaluation - what is measured, under what assumptions, and how results are interpreted. We define evaluation as the full set of processes, tools, datasets, benchmarks, and practices used to test, stress-test, audit, compare, and interpret AI/ML systems across their lifecycle. Datasets remain central to the track, both as components of evaluations and as resources used across the AI/ML lifecycle (training, fine-tuning, testing, auditing). However, dataset submissions should also clarify how they should be meaningfully used in evaluative practices rather than being endpoints in themselves. Subm
Call For Evaluations & Datasets 2026 CSP Test --> NeurIPS 2026 Evaluations & Datasets Track Call for Papers The NeurIPS Evaluations & Datasets (E&D) Track invites submissions that advance the science and practice of evaluation in AI/ML including the development and use of datasets and other resources. Formerly the Datasets & Benchmarks Track, the E&D Track reflects a shift and broadening in scope: evaluation becomes an object of scientific study in its own right . Scientific debates increasingly hinge on evaluation - what is measured, under what assumptions, and how results are interpreted. We
Explore this link on the map →related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- ACL 2026 Workshop on Evaluating Evaluations (EvalEval) | EvalEval Coalitionevalevalai.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- A statistical approach to model evaluations \ Anthropicanthropic.com
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- The bitter lesson of LLM evalsparsed.com
- A starter guide for evals — AI Alignment Forumalignmentforum.org
- Dataset list - A list of the biggest machine learning datasetsdatasetlist.com
- Preface - The Emerging Science of Machine Learning Benchmarksmlbenchmarks.org
- The Data Cards Playbook - Data Cards Playbooksites.research.google
- We Need A ‘Science of Evals’ – Apollo Researchapolloresearch.ai