Metrics of Agent Ability - METR
metr.org · 4,076 words · saved by 1 readers
Tom Cunningham surveys metrics for comparing AI agent capability as performance changes with expenditure and relative to human performance.
Metrics of Agent Ability - METR Our Work Research Notes Updates Risk Assessment About Donate Careers Search --> Our Work Research Notes Updates Risk Assessment About Donate Careers Menu × Research note: Metrics of Agent Ability CONTRIBUTORS Tom Cunningham DATE July 24, 2026 SHARE Copy Link Citation BibTeX Citation × @misc { metr-2026-metrics-of-model-ability , title = {Metrics of Agent Ability} , author = {Tom Cunningham} , howpublished = {\url{https://metr.org/notes/2026-07-24-metrics-of-model-ability/}} , year = {2026} , month = {07} , } Copy Tom Cunningham **Goals of this post**
saved by
related reading
- An Apple-Picking Model of AI R&D | Tom Cunningham – Tom Cunninghamtecunningham.github.io
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- More compute, more capability: Why AI agent evaluations need to account for test-time compute | AISI Workaisi.gov.uk
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPTmetr.org
- We spent 2 hours working in the future - METRmetr.org
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Economics and Transformative AI | Tom Cunningham – Tom Cunninghamtecunningham.github.io
- Humans Still Beat AI in the Long Horizon: Revisiting Test-Time Scaling in the Agent Era | Qiuyang Mangjoyemang33.github.io
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- After Automation | Everyevery.to
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com