✳flâneur — a map of the web's best reading
Training as we know it will end | Vintage Data
vintagedata.org · 1,860 words · saved by 1 readers
Old data, new models
Training as we know it will end | Vintage Data Training as we know it might end Pierre-Carl Langlais, May 8, 2025 This post was originally supposed to be about synthetic data. It is finally about a vibe shift. Over the last months, there is an accumulated amount of evidence that models are changing and synthetic data is roughly at the center of it. Too many developments have suddenly challenged firmly held assumptions: the rise of tiny reasoners, the incredible data efficiency of reinforcement learning, the rapid expansion of agentified training run into entire simulated environment. All this
Explore this link on the map →related reading
- Synthetic Pretraining | Vintage Datavintagedata.org
- DeepSeek-R1arxiv.org
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretrainingdatologyai.com
- Explore | alphaXivalphaxiv.org
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- Why reasoning models will generalize - by Nathan Lambertinterconnects.ai
- [2502.19402] General Reasoning Requires Learning to Reason from the Get-goar5iv.labs.arxiv.org
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Generative AI's Act o1: The Reasoning Era Begins | Sequoia Capitalsequoiacap.com