Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
arxiv.org · 4,873 words · saved by 1 readers
N/A
2025-7-4 Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding Chengyue Wu1,2* Hao Zhang2* Shuchen Xue4 Zhijian Liu2 Shizhe Diao2 Ligeng Zhu2 Ping Luo1 Song Han2,3 Enze Xie2 1 2 3 4 The University of Hong Kong…
related reading
- Accelerating Diffusion LLMs via Adaptive Parallel Decodingarxiv.org
- Fast-dLLM v2nvlabs.github.io
- Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusionarxiv.org
- Speculative Decoding - philkravphilkrav.com
- CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Creditsarxiv.org
- Large Language Diffusion Modelsarxiv.org
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- Fast Inference from Transformers via Speculative Decodingarxiv.org
- Esoteric Language Modelsarxiv.org
- Optimizing inference · Hugging Facehuggingface.co
- Kuleshov Group | How to Build a Diffusion Language Modelkuleshov-group.github.io