dr. jack morris on X: "some hypotheses for what “better pretraining” could mean - integration with other training stages: i’m guessing they’re finally at a point where post-training perf (eg SWE-Bench) can be used as signal for pretraining eng decisions - filtering: scaling approaches like influence" / X
To view keyboard shortcuts, press question mark View keyboard shortcuts Post See new posts Conversation dr. jack morris @jxmnop some hypotheses for what “better pretraining” could mean - integration with other training stages: i’m guessing they’re finally at a point where post-training perf (eg SWE-Bench) can be used as signal for pretraining eng decisions - filtering: scaling approaches like influence functions for getting rid of datapoints that don’t help eval perf - synthetic data: using rephrasing to upsample certain useful documents and make them more amenable to reasoning - mixing: more principled & scalable approaches for determining mixing coefficients - new data: purchasing and scanning more books, transcribing YouTube, buying private token collections like news articles - smart packing: there are various ways to group documents into batches that work better, especially for long-context stuff - systems: more data, more flops Quote Oriol Vinyals @OriolVinyalsML · Nov 18 Th
Jack Morris @jxmnop some hypotheses for what “better pretraining” could mean - integration with other training stages: i’m guessing they’re finally at a point where post-training perf (eg SWE-Bench) can be used as signal for pretraining eng decisions - filtering: scaling approaches like influence functions for getting rid of datapoints that don’t help eval perf - synthetic data: using rephrasing to upsample certain useful documents and make them more amenable to reasoning - mixing: more principled & scalable approaches for determining mixing coefficients - new data: purchasing and scanning mor
Explore this link on the map →related reading
- Elicitation, the simplest way to understand post-traininginterconnects.ai
- The Scaling Hypothesis · Gwern.netgwern.net
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- A Bitter Lesson for Data Filteringarxiv.org
- [2509.14786] Pre-training under infinite computearxiv.org
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Composer2.pdfcursor.com
- PostTrainBenchposttrainbench.com
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraintsarxiv.org
- BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretrainingdatologyai.com
- Pre, Mid, Post-Training Way of Life - by Tina Hefakepixels.substack.com