A Few Things I Learned About Evals - Ryan Bloom
I spent most of 2024 working on capability evaluations for AI agents. This post collects some salient takeaways from that work. Although much of what I learned was thanks to my kind and talented colleagues at METR, this post represents my own views only. A basic and general challenge with evaluation—one I expect is familiar to anyone who’s designed a standardized test or managed a work trial—is that an eval is not the real world, but it’s supposed to tell us something about the real world. Even with relatively flexible infrastructure, there are frequent trade-offs between practicality and validity. How do we quickly, cheaply tell if agents can manage long-running tasks on an expensive GPU cluster? What’s a controlled, reproducible way to assess how they deal with messy environments? At a high level, a good place to start is defining a budget and other practical constraints. Then take the overarching question (e.g. “does model X pose a catastrophic risk to society?”) and keep breaking i
I spent most of 2024 working on capability evaluations for AI agents. This post collects some salient takeaways from that work. Although much of what I learned was thanks to my kind and talented colleagues at METR, this post represents my own views only. To design a good eval, set the right constraints A basic and general challenge with evaluation—one I expect is familiar to anyone who's designed a standardized test or managed a work trial—is that an eval is not the real world, but it's supposed to tell us something about the real world. Even with relatively flexible infrastructure, there are
related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- More compute, more capability: Why AI agent evaluations need to account for test-time compute | AISI Workaisi.gov.uk
- AI agent evaluation frameworks for production - Vercelvercel.com
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- Successful language model evals - Jason Weijasonwei.net
- Agentic Evals Pyramidrwilinski.ai
- f316275b44ee2de533102913828a8107-Paper-Datasets_and_Benchmarks_Track.pdfproceedings.neurips.cc
- Things I learned at OpenAI - by Karina Nguyen - sémaphoresemaphore.substack.com