Evals as Theory Building — High Performance AI Lab
An eval is more than a score. It is a working theory of success—and it needs rigor, verification, and proof before its verdict is allowed to change a system.
“No one realized that the book and the labyrinth were one and the same.” Every organization has a sentence that can change the direction of its work: The numbers look good. Behind that sentence there is usually an evaluation—an eval—whether anybody calls it one or not. A pilot has been compared with the process it might replace. A supplier has been tested against another. A sample of customer interactions has been turned into a rating. Some part of reality has been selected, observed, and given a verdict. In Jorge Luis Borges’s one-paragraph story “On Exactitude in Science”, an empire…
saved by
related reading
- How AI evals are changing product managementmanialabs.substack.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- Successful language model evals - Jason Weijasonwei.net
- Ankur Goyal (@ankrgyl) on Xx.com
- A statistical approach to model evaluations \ Anthropicanthropic.com
- Agentic Evals Pyramidrwilinski.ai
- The bitter lesson of LLM evalsparsed.com
- AI agent evaluation frameworks for production - Vercelvercel.com
- A starter guide for evals — AI Alignment Forumalignmentforum.org
- Giovanni D'Antoniogiovannidantonio.com