flâneur — a map of the web's best reading

Linear Attention Fundamentals | Hailey Schoelkopf

haileyschoelkopf.github.io · 1,893 words · saved by 2 readers

The basics of linear attention in sub-quadratic language model architectures.

Linear Attention Fundamentals | Hailey Schoelkopf Linear Attention Fundamentals The basics of linear attention in sub-quadratic language model architectures. Introduction This post will be a short overview recapping key formulas and intuitions around the increasingly-popular family of methods under the umbrella of Linear Attention, first introduced by Katharopoulos et al. (2020) . The material in this post is also covered excellently by Yang, Wang et al. (2023) and Yang et al. (2024). This post will assume familiarity with the transformer architecture and softmax attention, and with KV caching

Explore this link on the map →

saved by

related reading