D2F: Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
Diffusion Large Language Models (dLLMs) have long held the promise of ultra-fast text generation through parallel decoding. Yet, this promise has remained unfulfilled—until now. In practice, open-source dLLMs have consistently lagged behind their autoregressive (AR) counterparts like LLaMA3 in inference speed, hindering their real-world adoption. Today, we are excited to introduce Discrete Diffusion Forcing (D2F), a simple yet powerful strategy that shatters this speed limit. D2F creates a novel AR-diffusion hybrid model that achieves up to a 2.5x speedup over leading AR models and a staggering 50x acceleration over vanilla dLLMs, all while maintaining comparable output quality. The potential of dLLMs lies in parallel decoding, but two fundamental issues have historically crippled their speed: D2F overcomes these bottlenecks by rethinking how dLLMs are trained and how they generate text. It's a two-part solution. At its core, D2F reframes generation as a block-autoregressive process. T
Diffusion Large Language Models (dLLMs) have long held the promise of ultra-fast text generation through parallel decoding. Yet, this promise has remained unfulfilled—until now. In practice, open-source dLLMs have consistently lagged behind their autoregressive (AR) counterparts like LLaMA3 in inference speed, hindering their real-world adoption. Today, we are excited to introduce Discrete Diffusion Forcing (D2F), a simple yet powerful strategy that shatters this speed limit. D2F creates a novel AR-diffusion hybrid model that achieves up to a 2.5x speedup over leading AR models and a staggerin
Explore this link on the map →