flâneur — a map of the web's best reading

dr. jack morris on X: "some hypotheses for what “better pretraining” could mean - integration with other training stages: i’m guessing they’re finally at a point where post-training perf (eg SWE-Bench) can be used as signal for pretraining eng decisions - filtering: scaling approaches like influence" / X

x.com · 208 words · saved by 1 readers

To view keyboard shortcuts, press question mark View keyboard shortcuts Post See new posts Conversation dr. jack morris @jxmnop some hypotheses for what “better pretraining” could mean - integration with other training stages: i’m guessing they’re finally at a point where post-training perf (eg SWE-Bench) can be used as signal for pretraining eng decisions - filtering: scaling approaches like influence functions for getting rid of datapoints that don’t help eval perf - synthetic data: using rephrasing to upsample certain useful documents and make them more amenable to reasoning - mixing: more principled & scalable approaches for determining mixing coefficients - new data: purchasing and scanning more books, transcribing YouTube, buying private token collections like news articles - smart packing: there are various ways to group documents into batches that work better, especially for long-context stuff - systems: more data, more flops Quote Oriol Vinyals @OriolVinyalsML · Nov 18 Th

Jack Morris @jxmnop some hypotheses for what “better pretraining” could mean - integration with other training stages: i’m guessing they’re finally at a point where post-training perf (eg SWE-Bench) can be used as signal for pretraining eng decisions - filtering: scaling approaches like influence functions for getting rid of datapoints that don’t help eval perf - synthetic data: using rephrasing to upsample certain useful documents and make them more amenable to reasoning - mixing: more principled & scalable approaches for determining mixing coefficients - new data: purchasing and scanning mor

Explore this link on the map →

related reading