flâneur

Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Models

arxiv.org · 5,903 words · saved by 1 readers

N/A

Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Models Steve Yadlowsky, Lyric Doshi, Nilesh Tripuraneni {yadlowsky, lyric, nileshtrip}@google.com Google DeepMind arXiv:2311.00871v1 [cs.LG] 1 Nov 2023 November 3,…

related reading