Data Distribution Shifts and Monitoring
Note: This note is a work-in-progress, created for the course CS 329S: Machine Learning Systems Design (Stanford, 2022). For the fully developed text, see the book Designing Machine Learning Systems (Chip Huyen, O’Reilly 2022). Slides (much shorter 😁). Original Google Docs version.
Note : This note is a work-in-progress, created for the course CS 329S: Machine Learning Systems Design (Stanford, 2022). For the fully developed text, see the book Designing Machine Learning Systems (Chip Huyen, O’Reilly 2022). Slides (much shorter 😁). Original Google Docs version . Let’s start the note with a story I was told by an executive that many readers might be able to relate to. About two years ago, his company hired a consulting firm to develop an ML model to help them predict how many of each grocery item they’d need next week, so they could restock the items accordingly. The cons
Explore this link on the map →related reading
- Real-time machine learning: challenges and solutionshuyenchip.com
- Why you need to improve your training data, and how to do it << Pete Warden's blogpetewarden.com
- Thinking about High-Quality Human Data | Lil'Loglilianweng.github.io
- deeplearningbook.org/contents/ml.htmldeeplearningbook.org
- A Course in Machine Learningciml.info
- Why are Machine Learning Projects so Hard to Manage? | by Lukas Biewald | Mediummedium.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- The bitter lesson of LLM evalsparsed.com
- DataRater: Meta-Learned Dataset Curationarxiv.org
- State of Data (Jan 2026)seancai.com
- Concept drift - Wikipediaen.wikipedia.org
- Preface - The Emerging Science of Machine Learning Benchmarksmlbenchmarks.org