Wide-In, Narrow-Out: Revokable Decoding for Efficient and Effective DLLMs
arxiv.org · 6,878 words · saved by 1 readers
N/A
Preprint W IDE -I N , NARROW-O UT: R EVOKABLE D ECODING FOR E FFICIENT AND E FFECTIVE DLLM S Feng Hong1,∗ Geng Yu1,∗ Yushi Ye1 Haicheng Huang1 Huangjie Zheng2 Ya Zhang3 Yanfeng Wang3 Jiangchao Yao1 1 Cooperative Medianet Innovation Center, Shanghai Jiao Tong University 2…
related reading
- Speculative Decoding - philkravphilkrav.com
- Speculative Decoding: How It Evolved, When It Stays Lossless, and What's Nextneurips2026-speculative-decoding.vercel.app
- Fast Inference from Transformers via Speculative Decodingarxiv.org
- How speculative decoding makes LLMs go brrr – Leonie Monigattileoniemonigatti.com
- Looking back at speculative decodingresearch.google
- Fast Inference from Transformers via Speculative Decodingarxiv.org
- Large Language Diffusion Modelsarxiv.org
- Accelerating Diffusion LLMs via Adaptive Parallel Decodingarxiv.org
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decodingarxiv.org
- Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusionarxiv.org
- CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Creditsarxiv.org
- Fast Inference from Transformers via Speculative Decodingarxiv.org