Thinking about High-Quality Human Data | Lil'Log
[Special thank you to Ian Kivlichan for many useful pointers (E.g. the 100+ year old Nature paper “Vox populi”) and nice feedback. 🙏 ] High-quality data is the fuel for modern data deep learning model training. Most of the task-specific labeled data comes from human annotation, such as classification task or RLHF labeling (which can be constructed as classification format) for LLM alignment training. Lots of ML techniques in the post can help with data quality, but fundamentally human data collection involves attention to details and careful execution. The community knows the value of high quality data, but somehow we have this subtle impression that “Everyone wants to do the model work, not the data work” (Sambasivan et al. 2021). Collecting human data involve a set of operation steps and every step contributes to the data quality: Vox populi (originally “Vox populi, vox Dei”), a Latin phrase, means the voice of people. A short paper named was the same name was published in 1907 on N
Table of Contents Human Raters ↔ Data Quality The Wisdom of the Crowd Rater Agreement Rater Disagreement & Two Paradigms Data Quality ↔ Model Training Influence Functions Prediction Changes during Training Noisy Cross-Validation Citation References [Special thank you to Ian Kivlichan for many useful pointers (E.g. the 100+ year old Nature paper "Vox populi") and nice feedback. 🙏 ] High-quality data is the fuel for modern data deep learning model training. Most of the task-specific labeled data comes from human annotation, such as classification task or RLHF labeling (which can be constructed
Explore this link on the map →saved by
related reading
- Why you need to improve your training data, and how to do it << Pete Warden's blogpetewarden.com
- DataRater: Meta-Learned Dataset Curationarxiv.org
- Good QC for RL Dataseancai.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- State of Data (Jan 2026)seancai.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Training LLMs On My Loved Ones' Data - by joyce c. chenjoinreboot.org
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Inside the AI Factory: The Humans That Make Tech Seem Humannymag.com
- Learning with not Enough Data Part 1: Semi-Supervised Learning | Lil'Loglilianweng.github.io
- Training Data: What Is It? All About Machine Learning Training Dataappen.com