flâneur — a map of the web's best reading

DFlash: Block Diffusion for Flash Speculative Decoding

z-lab.ai · 1,191 words · saved by 1 readers

TL;DR: In this work, we introduce DFlash, a method utilizing a lightweight block diffusion model for drafting in speculative decoding. This enables efficient and high-quality parallel drafting, pushing the limits of speculative decoding. DFlash achieves up to 6.17 × 6.17× lossless acceleration for Qwen3-8B, nearly 2.5 × 2.5× faster than the state-of-the-art speculative decoding method EAGLE-3, as shown in Figure 1. Huggingface Models: [Qwen3-4B-DFlash-b16] [Qwen3-8B-DFlash-b16] [Qwen3-Coder-30B-A3B-DFlash] Your browser does not support the video tag. Figure 1. Speedup comparison between DFlash, EAGLE-3 against Autoregressive Decoding. Overall, DFlash achieves more than 2.5× higher speedup than EAGLE-3. Note. Due to the lack of official EAGLE-3 checkpoints for Qwen3-4B and Qwen3-8B, we compare against RedHatAI/Qwen3-8B-speculator.eagle3 in this blog. This checkpoint is trained using the open-source speculators framework. DFlash is now supported on SGLang, enabling high-throughput spec

DFlash: Block Diffusion for Flash Speculative Decoding - Z Lab Z Lab Home Faculty Team Projects Home Faculty Team Projects DFlash: Block Diffusion for Flash Speculative Decoding Jian Chen , Yesheng Liang , Zhijian Liu Preprint Paper Code Models Tap to play Tap to play code]:bg-black/[0.06] [&_:not(pre)>code]:px-1.5 [&_:not(pre)>code]:py-0.5 [&_:not(pre)>code]:rounded [&_:not(pre)>code]:text-[0.8125rem] [&_a:hover]:text-red-600 [&_a]:no-underline [&_a]:text-blue-700 [&_blockquote]:border-black/15 [&_blockquote]:border-l-[3px] [&_blockquote]:italic [&_

Explore this link on the map →

saved by

related reading