2026 ACL Workshop on Evaluating AI in Practice | EvalEval Coalition
This workshop focuses on AI evaluation in practice, centering the tensions and collaborations between model developers and evaluation researchers and aims to surface practical insights from across the evaluation ecosystem.
Update: We are excited to meet you in San Diego! Save the date: on 03/July/2026 around 8pm, we will co-host a Social together with the GEM workshop . Both of our workshops will then be on the day after. 🗓️ Schedule 04/July/2026 Afternoon 👋 2:00 PM – 2:05 PM | Welcome and Introduction (5 mins) Jennifer Mickel , EvalEval Workshop Co-Chair 🎙️ 2:05 PM – 2:45 PM | Panel Presentation (30 mins panel + 10 mins Q&A) Moderator: Leshem Choshen, MIT, IBM Research, MIT-IBM Watson AI Lab This panel brings together model developers and evaluation researchers to examine how evaluations are designed, interp
Explore this link on the map →related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- AI in 2025: gestalt — LessWronglesswrong.com
- The bitter lesson of LLM evalsparsed.com
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- [2408.02565] Reasons to Doubt the Impact of AI Risk Evaluationsarxiv.org
- A statistical approach to model evaluations \ Anthropicanthropic.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- A starter guide for evals — AI Alignment Forumalignmentforum.org
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- [2506.13023] A Practical Guide for Evaluating LLMs and LLM-Reliant Systemsarxiv.org
- pdfopenreview.net
- A starter guide for evals — LessWronglesswrong.com