DFlash: Block Diffusion for Flash Speculative Decoding
TL;DR: In this work, we introduce DFlash, a method utilizing a lightweight block diffusion model for drafting in speculative decoding. This enables efficient and high-quality parallel drafting, pushing the limits of speculative decoding. DFlash achieves up to 6.17 × 6.17× lossless acceleration for Qwen3-8B, nearly 2.5 × 2.5× faster than the state-of-the-art speculative decoding method EAGLE-3, as shown in Figure 1. Huggingface Models: [Qwen3-4B-DFlash-b16] [Qwen3-8B-DFlash-b16] [Qwen3-Coder-30B-A3B-DFlash] Your browser does not support the video tag. Figure 1. Speedup comparison between DFlash, EAGLE-3 against Autoregressive Decoding. Overall, DFlash achieves more than 2.5× higher speedup than EAGLE-3. Note. Due to the lack of official EAGLE-3 checkpoints for Qwen3-4B and Qwen3-8B, we compare against RedHatAI/Qwen3-8B-speculator.eagle3 in this blog. This checkpoint is trained using the open-source speculators framework. DFlash is now supported on SGLang, enabling high-throughput spec
DFlash: Block Diffusion for Flash Speculative Decoding - Z Lab Z Lab Home Faculty Team Projects Home Faculty Team Projects DFlash: Block Diffusion for Flash Speculative Decoding Jian Chen , Yesheng Liang , Zhijian Liu Preprint Paper Code Models Tap to play Tap to play code]:bg-black/[0.06] [&_:not(pre)>code]:px-1.5 [&_:not(pre)>code]:py-0.5 [&_:not(pre)>code]:rounded [&_:not(pre)>code]:text-[0.8125rem] [&_a:hover]:text-red-600 [&_a]:no-underline [&_a]:text-blue-700 [&_blockquote]:border-black/15 [&_blockquote]:border-l-[3px] [&_blockquote]:italic [&_
Explore this link on the map →saved by
related reading
- Speculative Decoding - philkravphilkrav.com
- Fast Inference from Transformers via Speculative Decodingarxiv.org
- Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusionarxiv.org
- Speculative Decoding - Deep Dive — ROCm Blogsrocm.blogs.amd.com
- Speculative decodingaarnphm.xyz
- Nemotron_Diffusion_Tech_Report_v1.pdfd1qx31qr3h6wln.cloudfront.net
- Optimizing inference · Hugging Facehuggingface.co
- Large Language Diffusion Modelsarxiv.org
- Best practices to accelerate inference for large-scale production workloadstogether.ai
- Esoteric Language Modelsarxiv.org
- The economics of speculative decoding | Doublewordblog.doubleword.ai
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com