Giovanni D'Antonio
giovannidantonio.com · 3,862 words · saved by 1 readers
Giovanni D'Antonio explores intelligence and human-AI collaboration at Thinking Machines Lab.
Evals are suddenly everywhere: company strategy, political decisions, advertising. Yet there is surprisingly little material explaining the science behind them. This three-part series is my attempt to explain benchmarking from first principles. We will start with the basics: Why benchmark? What question is an evaluation supposed to answer? What should we benchmark? Which tasks, users, environments, and capabilities should the data represent? How should we benchmark? How do metrics, sampling, uncertainty, and statistics turn into a score? In Part II, we will go deeper into the ways…
saved by
related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- [1911.01547] On the Measure of Intelligencearxiv.org
- Successful language model evals - Jason Weijasonwei.net
- f316275b44ee2de533102913828a8107-Paper-Datasets_and_Benchmarks_Track.pdfproceedings.neurips.cc
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- On measuring AInikilravi.substack.com
- A statistical approach to model evaluations \ Anthropicanthropic.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Preface - The Emerging Science of Machine Learning Benchmarksmlbenchmarks.org
- AI’s capabilities may be exaggerated by flawed tests, study saysnbcnews.com
- Things I learned at OpenAI - by Karina Nguyen - sémaphoresemaphore.substack.com