Good QC for RL Data
seancai.com · 2,670 words · saved by 4 readers
Read my thoughts on Good QC for RL Data
This is a preview piece - for more writing and a state of data report, visit my substack . In January, I proposed a new definition for Type 1, Type 2 data, pending drastic need from the data industry on how to evaluate data quality. A conscious side-effect of the shift to longer horizon regiments is increased need for model-based QA, far beyond the body-shop capabilities of current day data companies. The progression of what data markets we entered first directly corresponded to how verifiable we could make each one. We filtered the hard domains out of the field at the infrastructure layer, fi
saved by
related reading
- State of Data (Jan 2026)seancai.com
- Vikram Aditya (@viks_rum) on Xx.com
- The Bitter Lesson - RL Environments Versionseancai.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Thinking about High-Quality Human Data | Lil'Loglilianweng.github.io
- Our Problems · Proximalproximal.ai
- RL Environments and RL for Science: Data Foundries and Multi-Agent Architecturesnewsletter.semianalysis.com
- Cheap RL tasks will waste compute | Mechanize, Inc.mechanize.work
- Reinforcement Learning With Verifiable Rewards: How Data and Verifiers Shape RLVRsnorkel.ai
- DataRater: Meta-Learned Dataset Curationarxiv.org
- We need 3rd party Training-Run Assessments — LessWronglesswrong.com
- A World of Verifiable Domainsseancai.com