flâneur

Giovanni D'Antonio

giovannidantonio.com · 3,862 words · saved by 1 readers

Giovanni D'Antonio explores intelligence and human-AI collaboration at Thinking Machines Lab.

Evals are suddenly everywhere: company strategy, political decisions, advertising. Yet there is surprisingly little material explaining the science behind them. This three-part series is my attempt to explain benchmarking from first principles. We will start with the basics: Why benchmark? What question is an evaluation supposed to answer? What should we benchmark? Which tasks, users, environments, and capabilities should the data represent? How should we benchmark? How do metrics, sampling, uncertainty, and statistics turn into a score? In Part II, we will go deeper into the ways…

saved by

related reading