flâneur — a map of the web's best reading

The Transformer Family Version 2.0 | Lil'Log

lilianweng.github.io · 9,727 words · saved by 1 readers

Many new Transformer architecture improvements have been proposed since my last post on “The Transformer Family” about three years ago. Here I did a big refactoring and enrichment of that 2020 post — restructure the hierarchy of sections and improve many sections with more recent papers. Version 2.0 is a superset of the old version, about twice the length. Notations Symbol Meaning $d$ The model size / hidden state dimension / positional encoding size.

Table of Contents Notations Transformer Basics Attention and Self-Attention Multi-Head Self-Attention Encoder-Decoder Architecture Positional Encoding Sinusoidal Positional Encoding Learned Positional Encoding Relative Position Encoding Rotary Position Embedding Longer Context Context Memory Non-Differentiable External Memory Distance-Enhanced Attention Scores Make it Recurrent Adaptive Modeling Adaptive Attention Span Depth-Adaptive Transformer Efficient Attention Sparse Attention Patterns Fixed Local Context Strided Context Combination of Local and Global Context Content-based Attention Low-

Explore this link on the map →

related reading