flâneur — a map of the web's best reading

[2510.03280] Training Optimal Large Diffusion Language Models

ar5iv.labs.arxiv.org · 12,885 words · saved by 1 readers

We introduce Quokka, the first systematic scaling law for diffusion language models (DLMs), encompassing both compute-constrained and data-constrained regimes, and studying the key modeling and optimization designs. Quokka is a good friend of Chinchilla and provides wider scopes. We hope the results would bring short-term practical guidance in DLMs training and long-term inspirations for the whole AI community. We summarize some takeaways below: • Compute-constrained law. With fixed FLOPs 𝐶 , the optimal parameters 𝑁 opt ∝ 𝐶 0.5 and data size 𝐷 opt ∝ 𝐶 0.5 , scaling at the same pace; DLMs are 2–5 × more data-hungry than autoregressive (AR) models at the same 𝐶 —favor smaller models and larger corpora. We provide a direct comparison with Chinchilla scaling law coefficients in Table 1 and their practical optimal allocation comparisons in Table 2. • Data-constrained law. Validation loss is U-shaped in epochs 𝑒 ; the onset of overfitting scales roughly as 𝑒 opt ∝ 𝑈 𝐷 0.39

\uselogo Training Optimal Large Diffusion Language Models Jinjie Ni 1 Correspondence to: Jinjie Ni < jinjieni@nus.edu.sg > Qian Liu Chao Du 2 Longxu Dou 2 Hang Yan 4 Zili Wang 3 Tianyu Pang 2 Michael Qizhe Shieh 1 1 National University of Singapore 2 Sea AI Lab 3 StepFun 4 Shanghai Qiji Zhifeng Co Ltd Abstract We introduce Quokka , the first systematic scaling law for diffusion language models (DLMs), encompassing both compute-constrained and data-constrained regimes, and studying the key modeling and optimization designs. Quokka is a good friend of Chinchilla and provides wider scopes. We hop

Explore this link on the map →

related reading