2302.01318
arxiv.org · 5,060 words · saved by 1 readers
N/A
2023-2-3 Accelerating Large Language Model Decoding with Speculative Sampling Charlie Chen1 , Sebastian Borgeaud1 , Geoffrey Irving1 , Jean-Baptiste Lespiau1 , Laurent Sifre1 and John Jumper1 1 All authors from DeepMind We present speculative sampling, an algorithm for accelerating transformer decoding by enabling the…
related reading
- Fast Inference from Transformers via Speculative Decodingarxiv.org
- Fast Inference from Transformers via Speculative Decodingarxiv.org
- Speculative Decoding - philkravphilkrav.com
- Fast Inference from Transformers via Speculative Decodingarxiv.org
- Looking back at speculative decodingresearch.google
- Speculative Decoding: How It Evolved, When It Stays Lossless, and What's Nextneurips2026-speculative-decoding.vercel.app
- How speculative decoding makes LLMs go brrr – Leonie Monigattileoniemonigatti.com
- [2402.12374] Sequoia: Scalable, Robust, and Hardware-aware Speculative Decodingarxiv.org
- Decoding Speculative Decoding from First Principlesjwlabs.vercel.app
- Speculative Speculative Decodingarxiv.org
- All About Transformer Inferencejax-ml.github.io
- Speculative decodingaarnphm.xyz