The Transformer Family Version 2.0 | Lil'Log
Many new Transformer architecture improvements have been proposed since my last post on “The Transformer Family” about three years ago. Here I did a big refactoring and enrichment of that 2020 post — restructure the hierarchy of sections and improve many sections with more recent papers. Version 2.0 is a superset of the old version, about twice the length. Notations Symbol Meaning $d$ The model size / hidden state dimension / positional encoding size.
Table of Contents Notations Transformer Basics Attention and Self-Attention Multi-Head Self-Attention Encoder-Decoder Architecture Positional Encoding Sinusoidal Positional Encoding Learned Positional Encoding Relative Position Encoding Rotary Position Embedding Longer Context Context Memory Non-Differentiable External Memory Distance-Enhanced Attention Scores Make it Recurrent Adaptive Modeling Adaptive Attention Span Depth-Adaptive Transformer Efficient Attention Sparse Attention Patterns Fixed Local Context Strided Context Combination of Local and Global Context Content-based Attention Low-
Explore this link on the map →related reading
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- transformer_attention.pdfarxiv.org
- Everything About Transformerskrupadave.com
- Transformers from Scratche2eml.school
- 1706.03762arxiv.org
- The Annotated Transformernlp.seas.harvard.edu
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Transformer (deep learning) - Wikipediaen.wikipedia.org
- Transformer (deep learning) - Wikipediaen.wikipedia.org
- Transformers from scratch | peterbloem.nlpeterbloem.nl
- The Annotated Transformernlp.seas.harvard.edu