Attention Normalizes the Wrong Norm | Convergent Thinking
We treat attention as a solved primitive. You can fuse kernels, tile memory access, quantize weights, but the formula itself is finished. It isn’t. Softmax normalizes the L1 norm to 1. Variance preservation requires the L2 norm to equal 1. These constraints differ. The mismatch causes attention output variance to collapse as sequence length grows, forcing models to learn position-dependent compensation. That compensation doesn’t transfer to unseen lengths. The fix is changing one norm. Attention computes ∑ 𝑖 𝐴 𝑖 𝑉 𝑖 . If the 𝑉 𝑖 have unit variance and are uncorrelated: Softmax constrains ‖ 𝐴 ‖ 1 = 1 , not ‖ 𝐴 ‖ 2 . For uniform attention over 𝑁 tokens, ‖ 𝐴 ‖ 2 = 1 / 𝑁 . Variance collapses as context grows. With L1 softmax, output magnitude depends on both sparsity and sequence length. With L2 softmax, it doesn’t. The same holds for gradients. Two identical models, differing only in softmax norm, trained to count occurrences of a target token. Training uses staged seq
We treat attention as a solved primitive. You can fuse kernels, tile memory access, quantize weights, but the formula itself is finished. It isn’t. Softmax normalizes the L1 norm to 1. Variance preservation requires the L2 norm to equal 1. These constraints differ. The mismatch causes attention output variance to collapse as sequence length grows, forcing models to learn position-dependent compensation. That compensation doesn’t transfer to unseen lengths. The fix is changing one norm. The Math Attention computes $\sum_i A_i V_i$. If the $V_i$ have unit variance and are uncorrelated:…
saved by
related reading
- Attention Is Off By One – Evan Millerevanmiller.org
- A short note on some aspects of long context attention | nor's blognor-blog.pages.dev
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Extending Context is Hard | kaiokendevkaiokendev.github.io
- Linear Attention Fundamentals | Hailey Schoelkopfhaileyschoelkopf.github.io
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- An Intuition for Attention | Jay Modyjaykmody.com
- Understanding Attention in LLMs | Bartosz Milewski's Programming Cafebartoszmilewski.com
- The Annotated Transformernlp.seas.harvard.edu
- Log-Linear Attentionarxiv.org
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Biao's Bloghebiao064.github.io