Preface - The Emerging Science of Machine Learning Benchmarks
Machine learning turns on one simple trick: Split your data into training and test sets. Anything goes on the training set; rank the models on the test set. Let the model builders compete. Call this a benchmark. Machine learning researchers cherish a good tradition of lamenting the shortcomings of machine learning benchmarks. Critics argue that static test sets and metrics promote narrow research objectives, stifling more creative scientific pursuits. Benchmarks also incentivize gaming the metrics, leading to inflated scores. Goodhart’s law cautions against competing over statistical measurements, but benchmarking ignores the warning. Over time, critics say, researchers overfit to benchmark datasets, building models that exploit artifacts. As a result, test set performance draws a skewed picture of model capabilities, deceiving us especially when comparing humans and machines. Add to this a slew of reasons why things don’t transfer from benchmarks to the real world. These scorching cri
Preface - The Emerging Science of Machine Learning Benchmarks Preface ← Index · Top ↑ Chapter Preface Overview Who is this book for? Acknowledgments Machine learning turns on one simple trick: Split your data into training and test sets. Anything goes on the training set; rank the models on the test set. Let the model builders compete. Call this a benchmark . Machine learning researchers cherish a good tradition of lamenting the shortcomings of machine learning benchmarks. Critics argue that static test sets and metrics promote narrow research objectives, stifling more creativ
Explore this link on the map →saved by
related reading
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- Questionable practices in machine learningarxiv.org
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Devising ML Metrics | CAISsafe.ai
- Competing in a data science contest without reading the datablog.mrtz.org
- Five things to keep in mind while reading biology ML papersowlposting.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- AI’s capabilities may be exaggerated by flawed tests, study saysnbcnews.com
- A statistical approach to model evaluations \ Anthropicanthropic.com
- Templates for machine learning research papers | Neel Guhaneelguha.github.io
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Good QC for RL Dataseancai.com