Full Stack Optimization of Transformer Inference: a Survey
arxiv.org · 4,916 words · saved by 1 readers
N/A
Full Stack Optimization of Transformer Inference: a Survey Sehoon Kim∗ Coleman Hooper∗ Thanakul Wattanawong sehoonkim@berkeley.edu chooper@berkeley.edu j.wat@berkeley.edu UC Berkeley UC Berkeley UC Berkeley…
related reading
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- How To Scale Your Modeljax-ml.github.io
- All About Transformer Inferencejax-ml.github.io
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- All the Transformer Math You Need to Know | How To Scale Your Modeljax-ml.github.io
- [2207.00032] DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scalearxiv.org
- Efficiently Scaling Transformer Inferencearxiv.org
- Accelerated Inference for Large Transformer Models Using NVIDIA Triton Inference Server | NVIDIA Technical Blogdeveloper.nvidia.com
- The Annotated Transformernlp.seas.harvard.edu
- Transformers from Scratche2eml.school