flâneur — a map of the web's best reading

All the Transformer Math You Need to Know | How To Scale Your Model

jax-ml.github.io · 4,517 words · saved by 2 readers

Here we'll do a quick review of the Transformer architecture, specifically how to calculate FLOPs, bytes, and other quantities of interest.

All the Transformer Math You Need to Know | How To Scale Your Model All the Transformer Math You Need to Know Part 4 of How To Scale Your Model ( Part 3: Sharding | Part 5: Training ) Here we'll do a quick review of the Transformer architecture, specifically how to calculate FLOPs, bytes, and other quantities of interest. Authors Affiliation Jacob Austin Google DeepMind Sholto Douglas Roy Frostig Anselm Levskaya Charlie Chen Sharad Vikram Federico Lebron Peter Choy Vinay Ramasesh Albert Webson Reiner Pope * Published Feb. 4, 2025 Counting Dots Let’s start with vectors \(x\), \(y\) and matrices

Explore this link on the map →

saved by

related reading