[2606.05405] Agents' Last Exam
Abstract:Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a benchmark designed to evaluate AI agents on long horizon, economically valuable, real world tasks with verifiable outcomes. Developed in collaboration with 250+ industry experts, ALE covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy). It is organized around a task taxonomy with 55 sub fields grouped into 13 industry clusters covering 1K+ tasks. Current results show that the hardest tier remains far from saturated: across mainstream harness and backbone configurations, the average full pass rate is below 1%. ALE is designed as a living benchmark: its task pool grows continuously as new workflows and industries are onboarded. More broadly, ALE is intended not merely as another leaderboard, but as an instrument for closing the gap between benchmark success and GDP relevant impact.
[2606.05405] Agents' Last Exam --> Computer Science > Artificial Intelligence arXiv:2606.05405 (cs) [Submitted on 3 Jun 2026 ( v1 ), last revised 11 Jun 2026 (this version, v2)] Title: Agents' Last Exam Authors: Yiyou Sun , Xinyang Han , Weichen Zhang , Yuanbo Pang , Tianyu Wang , Yuhan Cao , Yixiao Huang , Chris Duroiu , Haoyun Zhang , Jeffrey Lin , Weishu Zhang , Tyler Zeng , Ying Yan , Bo Liu , Hanson Wen , Mingyang Xu , Xiaoyuan Liu , Zimeng Chen , Weiyan Shi , Amanda Dsouza , Vincent Sunn Chen , Patrick Bryant , Carl Boettiger , Yamini Rangan , Bradley Rothenberg , Kyle Steinfeld , Arvind
Explore this link on the map →saved by
related reading
- PostTrainBenchposttrainbench.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- Measuring the performance of our models on real-world tasks | OpenAIopenai.com
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- After Automation | Everyevery.to
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering | alphaXivalphaxiv.org
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHubgithub.com
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com