Ankur Goyal on X: "Evals ≠ tests" / X
x.com · 807 words · saved by 1 readers
Evals ≠ tests
While building Braintrust I’ve been thinking about what evals are good at, what they should be used for, and how they differ from other parts of the software development lifecycle, like tests. Though it’s tempting to conflate evals and tests, doing so confuses the differences between them. Here’s how we think about it at Braintrust: A test is about if the thing can work An eval is about what the thing can and cannot do A mature eval practice is about what the thing should be doing, how often, and in what circumstances The last one is important when building agents. You need to test…
saved by
related reading
- Agentic Evals Pyramidrwilinski.ai
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Killing Coding Agent Slop With Adversarial Self-Playusetelos.ai
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- Agent behavioragentbehavior.dev
- Notes on the Software Factorybenedict.dev
- Senior SWE-Benchsenior-swe-bench.snorkel.ai
- Evals as Theory Building — High Performance AI Labhighperformanceailab.com
- AI agent evaluation frameworks for production - Vercelvercel.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- How to Eval AI Agents — The 2026 Guidehowtoeval.com
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com