flâneur — a map of the web's best reading

Preface - The Emerging Science of Machine Learning Benchmarks

mlbenchmarks.org · 2,713 words · saved by 1 readers

Machine learning turns on one simple trick: Split your data into training and test sets. Anything goes on the training set; rank the models on the test set. Let the model builders compete. Call this a benchmark. Machine learning researchers cherish a good tradition of lamenting the shortcomings of machine learning benchmarks. Critics argue that static test sets and metrics promote narrow research objectives, stifling more creative scientific pursuits. Benchmarks also incentivize gaming the metrics, leading to inflated scores. Goodhart’s law cautions against competing over statistical measurements, but benchmarking ignores the warning. Over time, critics say, researchers overfit to benchmark datasets, building models that exploit artifacts. As a result, test set performance draws a skewed picture of model capabilities, deceiving us especially when comparing humans and machines. Add to this a slew of reasons why things don’t transfer from benchmarks to the real world. These scorching cri

Preface - The Emerging Science of Machine Learning Benchmarks Preface ← Index · Top ↑ Chapter Preface Overview Who is this book for? Acknowledgments Machine learning turns on one simple trick: Split your data into training and test sets. Anything goes on the training set; rank the models on the test set. Let the model builders compete. Call this a benchmark . Machine learning researchers cherish a good tradition of lamenting the shortcomings of machine learning benchmarks. Critics argue that static test sets and metrics promote narrow research objectives, stifling more creativ

Explore this link on the map →

saved by

related reading