[2510.03280] Training Optimal Large Diffusion Language Models
We introduce Quokka, the first systematic scaling law for diffusion language models (DLMs), encompassing both compute-constrained and data-constrained regimes, and studying the key modeling and optimization designs. Quokka is a good friend of Chinchilla and provides wider scopes. We hope the results would bring short-term practical guidance in DLMs training and long-term inspirations for the whole AI community. We summarize some takeaways below: • Compute-constrained law. With fixed FLOPs 𝐶 , the optimal parameters 𝑁 opt ∝ 𝐶 0.5 and data size 𝐷 opt ∝ 𝐶 0.5 , scaling at the same pace; DLMs are 2–5 × more data-hungry than autoregressive (AR) models at the same 𝐶 —favor smaller models and larger corpora. We provide a direct comparison with Chinchilla scaling law coefficients in Table 1 and their practical optimal allocation comparisons in Table 2. • Data-constrained law. Validation loss is U-shaped in epochs 𝑒 ; the onset of overfitting scales roughly as 𝑒 opt ∝ 𝑈 𝐷 0.39
\uselogo Training Optimal Large Diffusion Language Models Jinjie Ni 1 Correspondence to: Jinjie Ni < jinjieni@nus.edu.sg > Qian Liu Chao Du 2 Longxu Dou 2 Hang Yan 4 Zili Wang 3 Tianyu Pang 2 Michael Qizhe Shieh 1 1 National University of Singapore 2 Sea AI Lab 3 StepFun 4 Shanghai Qiji Zhifeng Co Ltd Abstract We introduce Quokka , the first systematic scaling law for diffusion language models (DLMs), encompassing both compute-constrained and data-constrained regimes, and studying the key modeling and optimization designs. Quokka is a good friend of Chinchilla and provides wider scopes. We hop
Explore this link on the map →related reading
- Large Language Diffusion Modelsarxiv.org
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- Large Language Diffusion Modelsarxiv.org
- On neural scaling and the quanta hypothesisericjmichaud.com
- Esoteric Language Modelsarxiv.org
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraintsarxiv.org