2026 ACL Workshop on Evaluating AI in Practice | EvalEval Coalition
This workshop focuses on AI evaluation in practice, centering the tensions and collaborations between model developers and evaluation researchers and aims to surface practical insights from across the evaluation ecosystem.
Update: We are excited to meet you in San Diego! Save the date: on 03/July/2026 around 8pm, we will co-host a Social together with the GEM workshop . Both of our workshops will then be on the day after. 🗓️ Schedule 04/July/2026 Afternoon 👋 2:00 PM – 2:05 PM | Welcome and Introduction (5 mins) Jennifer Mickel , EvalEval Workshop Co-Chair 🎙️ 2:05 PM – 2:45 PM | Panel Presentation (30 mins panel + 10 mins Q&A) Moderator: Leshem Choshen, MIT, IBM Research, MIT-IBM Watson AI Lab This panel brings together model developers and evaluation researchers to examine how evaluations are designed, interp
related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Successful language model evals - Jason Weijasonwei.net
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- The bitter lesson of LLM evalsparsed.com
- A statistical approach to model evaluations \ Anthropicanthropic.com
- Things I learned at OpenAI - by Karina Nguyen - sémaphoresemaphore.substack.com
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- [2408.02565] Reasons to Doubt the Impact of AI Risk Evaluationsarxiv.org
- A starter guide for evals — AI Alignment Forumalignmentforum.org
- withhumans.pdfgleech.org
- Toward A Public Science of Model Behavior | Transluce AItransluce.org