flâneur

2302.01318

arxiv.org · 5,060 words · saved by 1 readers

N/A

2023-2-3 Accelerating Large Language Model Decoding with Speculative Sampling Charlie Chen1 , Sebastian Borgeaud1 , Geoffrey Irving1 , Jean-Baptiste Lespiau1 , Laurent Sifre1 and John Jumper1 1 All authors from DeepMind We present speculative sampling, an algorithm for accelerating transformer decoding by enabling the…

related reading