✳flâneur — a map of the web's best reading
Linear Attention Fundamentals | Hailey Schoelkopf
haileyschoelkopf.github.io · 1,893 words · saved by 2 readers
The basics of linear attention in sub-quadratic language model architectures.
Linear Attention Fundamentals | Hailey Schoelkopf Linear Attention Fundamentals The basics of linear attention in sub-quadratic language model architectures. Introduction This post will be a short overview recapping key formulas and intuitions around the increasingly-popular family of methods under the umbrella of Linear Attention, first introduced by Katharopoulos et al. (2020) . The material in this post is also covered excellently by Yang, Wang et al. (2023) and Yang et al. (2024). This post will assume familiarity with the transformer architecture and softmax attention, and with KV caching
Explore this link on the map →saved by
related reading
- DeltaNet Explained (Part I) | Songlin Yangsustcsonglin.github.io
- A short note on some aspects of long context attention | nor's blognor-blog.pages.dev
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Linear Attention Is All You Need | Towards Data Sciencetowardsdatascience.com
- Parallelizing Linear Transformers with the Delta Rule over Sequence Lengtharxiv.org
- Linear Transformers Are Faster After All – Manifest AImanifestai.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Overleaf Examplearxiv.org
- transformer_attention.pdfarxiv.org
- Mamba: The Easy Wayjackcook.com
- 1706.03762arxiv.org