✳flâneur — a map of the web's best reading
Synthetic Pretraining | Vintage Data
vintagedata.org · 4,034 words · saved by 1 readers
Old data, new models
Synthetic Pretraining | Vintage Data Synthetic pretraining Pierre-Carl Langlais, February 1, 2026 Pretraining data infrastructure used to be the most conservative part of a fast-moving AI world. Since GPT-3 we have been mostly scaling the usual mix of web crawls peppered with a few more select sources (including, controversially, digitized books). This is finally changing. In 2025, several major releases used extensive synthetic datasets before mid-training happens: Minimax , Trinity , K2 /K2.5, Nemotron-3 and, more speculatively, GPT-OSS . At Pleias we even experimented with full synthetic tr
Explore this link on the map →related reading
- BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretrainingdatologyai.com
- Training as we know it will end | Vintage Datavintagedata.org
- How to Generate and Use Synthetic Data for Finetuningeugeneyan.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- [2502.19402] General Reasoning Requires Learning to Reason from the Get-goar5iv.labs.arxiv.org
- GenAI Handbookgenai-handbook.github.io
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Alignment pretraining could backfire — LessWronglesswrong.com
- Machine Learning for Synthetic Data Generation: A Reviewarxiv.org