Nemotron_Diffusion_Tech_Report_v1.pdf
d1qx31qr3h6wln.cloudfront.net · 14,335 words · saved by 1 readers
N/A
# link_nesuh0z3tc.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=true - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - CreationDate=D:20260519170612Z - Creator=LaTeX with hyperref - ModDate=D:20260519170612Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.27 (TeX Live 2025) kpathsea version 6.4.1 - Producer=pdfTeX-1.40.27 - Trapped=False ## Contents ### Page 1 2026-5-19Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation DecodingYonggan Fu, Lexington Wh
saved by
related reading
- Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusionarxiv.org
- Large Language Diffusion Modelsarxiv.org
- Esoteric Language Modelsarxiv.org
- 2503.09573arxiv.org
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- Kuleshov Group | How to Build a Diffusion Language Modelkuleshov-group.github.io
- 2406.07524arxiv.org
- Speculative Decoding - philkravphilkrav.com
- 2409.02908arxiv.org
- Speculative Decoding: How It Evolved, When It Stays Lossless, and What's Nextneurips2026-speculative-decoding.vercel.app
- What are Diffusion Models?lilianweng.github.io
- Accelerating Diffusion LLMs via Adaptive Parallel Decodingarxiv.org