DFlash: Block Diffusion for Flash Speculative Decoding
TL;DR: In this work, we introduce DFlash, a method utilizing a lightweight block diffusion model for drafting in speculative decoding. This enables efficient and high-quality parallel drafting, pushing the limits of speculative decoding. DFlash achieves up to 6.17 × 6.17× lossless acceleration for Qwen3-8B, nearly 2.5 × 2.5× faster than the state-of-the-art speculative decoding method EAGLE-3, as shown in Figure 1. Huggingface Models: [Qwen3-4B-DFlash-b16] [Qwen3-8B-DFlash-b16] [Qwen3-Coder-30B-A3B-DFlash] Your browser does not support the video tag. Figure 1. Speedup comparison between DFlash, EAGLE-3 against Autoregressive Decoding. Overall, DFlash achieves more than 2.5× higher speedup than EAGLE-3. Note. Due to the lack of official EAGLE-3 checkpoints for Qwen3-4B and Qwen3-8B, we compare against RedHatAI/Qwen3-8B-speculator.eagle3 in this blog. This checkpoint is trained using the open-source speculators framework. DFlash is now supported on SGLang, enabling high-throughput spec
DFlash: Block Diffusion for Flash Speculative Decoding - Z Lab Z Lab Home Faculty Team Projects Home Faculty Team Projects DFlash: Block Diffusion for Flash Speculative Decoding Jian Chen , Yesheng Liang , Zhijian Liu Preprint Paper Code Models Tap to play Tap to play code]:bg-black/[0.06] [&_:not(pre)>code]:px-1.5 [&_:not(pre)>code]:py-0.5 [&_:not(pre)>code]:rounded [&_:not(pre)>code]:text-[0.8125rem] [&_a:hover]:text-red-600 [&_a]:no-underline [&_a]:text-blue-700 [&_blockquote]:border-black/15 [&_blockquote]:border-l-[3px] [&_blockquote]:italic [&_
saved by
related reading
- Speculative Decoding - philkravphilkrav.com
- Speculative Decoding: How It Evolved, When It Stays Lossless, and What's Nextneurips2026-speculative-decoding.vercel.app
- How speculative decoding makes LLMs go brrr – Leonie Monigattileoniemonigatti.com
- Fast Inference from Transformers via Speculative Decodingarxiv.org
- Speculative Speculative Decodingarxiv.org
- Fast Inference from Transformers via Speculative Decodingarxiv.org
- Looking back at speculative decodingresearch.google
- Decoding Speculative Decoding from First Principlesjwlabs.vercel.app
- [2402.12374] Sequoia: Scalable, Robust, and Hardware-aware Speculative Decodingarxiv.org
- Fast Inference from Transformers via Speculative Decodingarxiv.org
- Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusionarxiv.org
- Speculative Decoding - Deep Dive — ROCm Blogsrocm.blogs.amd.com