Measuring the performance of our models on real-world tasks | OpenAI
We’re introducing GDPval, a new evaluation that measures model performance on economically valuable, real-world tasks across 44 occupations. Our mission is to ensure that artificial general intelligence benefits all of humanity. As part of our mission, we want to transparently communicate progress on how AI models can help people in the real world. That’s why we’re introducing GDPval: a new evaluation designed to help us track how well our models and others perform on economically valuable, real-world tasks. We call this evaluation GDPval because we started with the concept of Gross Domestic Product (GDP) as a key economic indicator and drew tasks from the key occupations in the industries that contribute most to GDP. People often speculate about AI’s broader impact on society, but the clearest way to understand its potential is by looking at what models are already capable of doing. History shows that major technologies—from the internet to smartphones—took more than a decade to go fr
September 25, 2025 Publication Research Measuring the performance of our models on real-world tasks We’re introducing GDPval, a new evaluation that measures model performance on economically valuable, real-world tasks across 44 occupations. Read the paper (opens in a new window) Visit evals.openai.com (opens in a new window) Share Our mission is to ensure that artificial general intelligence benefits all of humanity. As part of our mission, we want to transparently communicate progress on how AI models can help people in the real world. That’s why we’re introducing GDPval: a new evaluation des
Explore this link on the map →saved by
related reading
- Economics and Transformative AI | Tom Cunningham – Tom Cunninghamtecunningham.github.io
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- After Automation | Everyevery.to
- gpt-4.pdfcdn.openai.com
- Introducing the Anthropic Economic Index \ Anthropicanthropic.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- GPT-4openai.com
- [2606.05405] Agents' Last Examarxiv.org
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- Labor market impacts of AI: A new measure and early evidence \ Anthropicanthropic.com
- AI Agent Benchmark for Real-World Professional Workflowsagents-last-exam.org