Yudhister Kumar
[1] Diffusion models seem to outperform traditional autoregressive models in the large data limit on token-prediction tasks. 1 Autoregressive models are still superior in the low-data/compute-limited regime, and the threshold at which diffusion models become optimal follows a power-law in the dataset size (typically exceeding the Chinchilla threshold by a large margin).2 Diffusion models also see performance gains under “trivial” data augmentation methods for far longer than autoregressive models (e.g. reordering tokens), and this is plausibly because the generation method is fundamentally non-causal? (Much of the performance gap can be recovered by implementing similar data augmentation methods in the AR case, but it’s unclear if this scales to tasks that require “cognition” in the human sense of the word). Not entirely clear how this translates to better performance on real-world tasks in the data-limited regime; it could be that the compute scaling necessary is simply prohibitive, a
[1] Diffusion models seem to outperform traditional autoregressive models in the large data limit on token-prediction tasks. 1 Autoregressive models are still superior in the low-data/compute-limited regime, and the threshold at which diffusion models become optimal follows a power-law in the dataset size (typically exceeding the Chinchilla threshold by a large margin). 2 Diffusion models also see performance gains under “trivial” data augmentation methods for far longer than autoregressive models (e.g. reordering tokens), and this is plausibly because the generation method is fundamentally no
related reading
- Diffusion Beats Autoregressive in Data-Constrained Settingsarxiv.org
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- What are Diffusion Models?lilianweng.github.io
- Diffusion is spectral autoregression – Sander Dielemansander.ai
- Large Language Diffusion Modelsarxiv.org
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- Diffusion language models – Sander Dielemansander.ai
- 2503.09573arxiv.org
- Diffusion is not necessarily Spectral Autoregression | Fabian Falckfabianfalck.com
- Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusionarxiv.org
- Kuleshov Group | How to Build a Diffusion Language Modelkuleshov-group.github.io
- Scalable Diffusion Models with Transformersopenaccess.thecvf.com