The Case for Evaluating Model Behaviors — AI Alignment Forum
Most evaluations of AI systems focus on their capabilities: how good they are at coding tasks, how effectively they can answer complex scientific questions, and so on. From a safety perspective, capability evaluations have a place: by understanding how close we are to different capabilities, and the rate of progress on them, we can forecast when different risks are likely to occur, as well as the broad shape of AI development. These capability evaluations were very useful to me when writing GPT-2030, and more recently I've found the METR time horizon graph useful for extrapolating the likely degree of autonomy of future agents. However, these evaluations also have pretty significant externalities: accurate capability measurements speed up capability research, and the work needed to fully elicit model capabilities involves developing agent scaffolds and other artifacts that directly advance model capabilities. This also means that AI labs are already highly incentivized to produce such
x The Case for Evaluating Model Behaviors — AI Alignment Forum AI Frontpage 16 The Case for Evaluating Model Behaviors by jsteinhardt 20th May 2026 4 min read 3 16 Most evaluations of AI systems focus on their capabilities: how good they are at coding tasks, how effectively they can answer complex scientific questions, and so on. From a safety perspective, capability evaluations have a place: by understanding how close we are to different capabilities, and the rate of progress on them, we can forecast when different risks are likely to occur, as well as the broad shape of AI development. These
Explore this link on the map →related reading
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Toward A Public Science of Model Behavior | Transluce AItransluce.org
- Chapter 3: LLM Evaluations - ARENAlearn.arena.education
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- [2408.02565] Reasons to Doubt the Impact of AI Risk Evaluationsarxiv.org
- Model evals for dangerous capabilities — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Off Target | CNAScnas.org
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- Realistic Evaluations Will Not Prevent Evaluation Awareness — LessWronglesswrong.com