Call For Evaluations & Datasets 2026
The NeurIPS Evaluations & Datasets (E&D) Track invites submissions that advance the science and practice of evaluation in AI/ML including the development and use of datasets and other resources. Formerly the Datasets & Benchmarks Track, the E&D Track reflects a shift and broadening in scope: evaluation becomes an object of scientific study in its own right. Scientific debates increasingly hinge on evaluation - what is measured, under what assumptions, and how results are interpreted. We define evaluation as the full set of processes, tools, datasets, benchmarks, and practices used to test, stress-test, audit, compare, and interpret AI/ML systems across their lifecycle. Datasets remain central to the track, both as components of evaluations and as resources used across the AI/ML lifecycle (training, fine-tuning, testing, auditing). However, dataset submissions should also clarify how they should be meaningfully used in evaluative practices rather than being endpoints in themselves. Subm
Call For Evaluations & Datasets 2026 CSP Test --> NeurIPS 2026 Evaluations & Datasets Track Call for Papers The NeurIPS Evaluations & Datasets (E&D) Track invites submissions that advance the science and practice of evaluation in AI/ML including the development and use of datasets and other resources. Formerly the Datasets & Benchmarks Track, the E&D Track reflects a shift and broadening in scope: evaluation becomes an object of scientific study in its own right . Scientific debates increasingly hinge on evaluation - what is measured, under what assumptions, and how results are interpreted. We
related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Successful language model evals - Jason Weijasonwei.net
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- ACL 2026 Workshop on Evaluating Evaluations (EvalEval) | EvalEval Coalitionevalevalai.com
- A statistical approach to model evaluations \ Anthropicanthropic.com
- CodaLab Worksheetsworksheets.codalab.org
- Toward A Public Science of Model Behavior | Transluce AItransluce.org
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Giovanni D'Antoniogiovannidantonio.com
- Scaling Laws Do Not Scalearxiv.org
- A starter guide for evals — AI Alignment Forumalignmentforum.org