flâneur

Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding

arxiv.org · 4,873 words · saved by 1 readers

N/A

2025-7-4 Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding Chengyue Wu1,2* Hao Zhang2* Shuchen Xue4 Zhijian Liu2 Shizhe Diao2 Ligeng Zhu2 Ping Luo1 Song Han2,3 Enze Xie2 1 2 3 4 The University of Hong Kong…

related reading