Measuring AI Ability to Complete Long Tasks - METR
We propose measuring AI performance in terms of the *length* of tasks AI agents can complete. We show that this metric has been consistently exponentially increasing over the past 6 years, with a doubling time of around 7 months. Extrapolating this trend predicts that, in under a decade, we will see AI agents that can independently complete a large fraction of software tasks that currently take humans days or weeks.
Measuring AI Ability to Complete Long Tasks - METR Our Work Research Notes Updates Risk Assessment About Donate Careers Search --> Our Work Research Notes Updates Risk Assessment About Donate Careers Menu × Measuring AI Ability to Complete Long Tasks We propose measuring AI performance in terms of the length of tasks AI agents can complete. We show that this metric has been consistently exponentially increasing over the past 6 years, with a doubling time of around 7 months. Extrapolating this trend predicts that, in under a decade, we will see AI agents that can independently complete a
Explore this link on the map →saved by
related reading
- I underestimated AI capabilities (again) - by Ajeya Cotraplanned-obsolescence.org
- A (Slightly) Mechanistic Theory for Exponentially Increasing AI Time Horizons? — LessWronglesswrong.com
- Clarifying and predicting AGI — LessWronglesswrong.com
- AI 2027ai-2027.com
- AI progress is about to speed up | Epoch AIepoch.ai
- We spent 2 hours working in the future - METRmetr.org
- Is there a Half-Life for the Success Rates of AI Agents? - Toby Ordtobyord.com
- How Does Time Horizon Vary Across Domains? - METRmetr.org
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- My picture of the present in AI — LessWronglesswrong.com
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity - METRmetr.org
- AIs can now often do massive easy-to-verify SWE tasks and I've updated towards shorter timelines — LessWronglesswrong.com