All the Transformer Math You Need to Know | How To Scale Your Model
jax-ml.github.io · 4,517 words · saved by 3 readers
Here we'll do a quick review of the Transformer architecture, specifically how to calculate FLOPs, bytes, and other quantities of interest.
All the Transformer Math You Need to Know | How To Scale Your Model All the Transformer Math You Need to Know Part 4 of How To Scale Your Model ( Part 3: Sharding | Part 5: Training ) Here we'll do a quick review of the Transformer architecture, specifically how to calculate FLOPs, bytes, and other quantities of interest. Authors Affiliation Jacob Austin Google DeepMind Sholto Douglas Roy Frostig Anselm Levskaya Charlie Chen Sharad Vikram Federico Lebron Peter Choy Vinay Ramasesh Albert Webson Reiner Pope * Published Feb. 4, 2025 Counting Dots Let’s start with vectors \(x\), \(y\) and matrices
saved by
related reading
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- The Annotated Transformernlp.seas.harvard.edu
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- Inside the Transformer: The Life of a Token - Aleksa Gordićaleksagordic.com
- Transformers from Scratche2eml.school
- Thinking like Transformersrush.github.io
- The Annotated Transformernlp.seas.harvard.edu
- Transformers from scratch | peterbloem.nlpeterbloem.nl
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org