DPad: Efficient Diffusion Language Models with Suffix Dropout
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. Diffusion-based Large Language Models (dLLMs) parallelize text generation by framing decoding as a denoising process, but suffer from high computational overhead since they predict all future suffix tokens at each step while retaining only a small fraction. We propose Diffusion Scratchpad (DPad), a training-free method that restricts attention to a small set of nearby suffix
DPad: Efficient Diffusion Language Models with Suffix Dropout Xinhua Chen 1 1 1 1 Equal contribution Sitao Huang 1 1 1 footnotemark: 1 Cong Guo 1 1 1 footnotemark: 1 2 2 2 Corresponding author: Cong Guo (cong.guo@duke.edu) Chiyue Wei 1 Yintao He 1 Jianyi Zhang 1 Hai “Hellen” Li 1 Yiran Chen 1 1 Duke University Abstract Diffusion-based Large Language Models (dLLMs) parallelize text generation by framing decoding as a denoising process, but suffer from high computational overhead since they predict all future suffix tokens at each step while retaining only a small fraction. We propose Diffusion
related reading
- Large Language Diffusion Modelsarxiv.org
- Esoteric Language Modelsarxiv.org
- 2409.02908arxiv.org
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decodingarxiv.org
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusionarxiv.org
- TiDAR: Think in Diffusion, Talk in Autoregressionalphaxiv.org
- Speculative Decoding - philkravphilkrav.com
- Kuleshov Group | How to Build a Diffusion Language Modelkuleshov-group.github.io
- Accelerating Diffusion LLMs via Adaptive Parallel Decodingarxiv.org
- 2503.09573arxiv.org
- Speculative Decoding: How It Evolved, When It Stays Lossless, and What's Nextneurips2026-speculative-decoding.vercel.app