✳flâneur — a map of the web's best reading
Metrics of Agent Ability - METR
metr.org · 4,076 words · saved by 1 readers
Tom Cunningham surveys metrics for comparing AI agent capability as performance changes with expenditure and relative to human performance.
Metrics of Agent Ability - METR Our Work Research Notes Updates Risk Assessment About Donate Careers Search --> Our Work Research Notes Updates Risk Assessment About Donate Careers Menu × Research note: Metrics of Agent Ability CONTRIBUTORS Tom Cunningham DATE July 24, 2026 SHARE Copy Link Citation BibTeX Citation × @misc { metr-2026-metrics-of-model-ability , title = {Metrics of Agent Ability} , author = {Tom Cunningham} , howpublished = {\url{https://metr.org/notes/2026-07-24-metrics-of-model-ability/}} , year = {2026} , month = {07} , } Copy Tom Cunningham **Goals of this post**
Explore this link on the map →related reading
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- Demystifying evals for AI agents \ Anthropicanthropic.com
- An Apple-Picking Model of AI R&D | Tom Cunningham – Tom Cunninghamtecunningham.github.io
- We spent 2 hours working in the future - METRmetr.org
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Economics and Transformative AI | Tom Cunningham – Tom Cunninghamtecunningham.github.io
- Humans Still Beat AI in the Long Horizon: Revisiting Test-Time Scaling in the Agent Era | Qiuyang Mangjoyemang33.github.io
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- A (Slightly) Mechanistic Theory for Exponentially Increasing AI Time Horizons? — LessWronglesswrong.com
- After Automation | Everyevery.to
- Is there a Half-Life for the Success Rates of AI Agents? - Toby Ordtobyord.com