Efficiently Scaling Transformer Inference
arxiv.org · 5,072 words · saved by 1 readers
N/A
E FFICIENTLY S CALING T RANSFORMER I NFERENCE Reiner Pope 1 Sholto Douglas 1 Aakanksha Chowdhery 1 Jacob Devlin 1 James Bradbury 1 Anselm Levskaya 1 Jonathan Heek 1 Kefan Xiao 1 Shivani Agrawal 1 Jeff Dean 1 A BSTRACT We study the problem of efficient generative inference for Transformer models, in one of its most challenging…
related reading
- All About Transformer Inferencejax-ml.github.io
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- How To Scale Your Modeljax-ml.github.io
- [2207.00032] DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scalearxiv.org
- Accelerated Inference for Large Transformer Models Using NVIDIA Triton Inference Server | NVIDIA Technical Blogdeveloper.nvidia.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Full Stack Optimization of Transformer Inference: a Surveyarxiv.org
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- All the Transformer Math You Need to Know | How To Scale Your Modeljax-ml.github.io
- Transformer inference tricks - by Finbarr Timbersartfintel.com