flâneur — a map of the web's best reading

Attention Normalizes the Wrong Norm | Convergent Thinking

convergentthinking.sh · 304 words · saved by 1 readers

We treat attention as a solved primitive. You can fuse kernels, tile memory access, quantize weights, but the formula itself is finished. It isn’t. Softmax normalizes the L1 norm to 1. Variance preservation requires the L2 norm to equal 1. These constraints differ. The mismatch causes attention output variance to collapse as sequence length grows, forcing models to learn position-dependent compensation. That compensation doesn’t transfer to unseen lengths. The fix is changing one norm. Attention computes ∑ 𝑖 𝐴 𝑖 𝑉 𝑖 . If the 𝑉 𝑖 have unit variance and are uncorrelated: Softmax constrains ‖ 𝐴 ‖ 1 = 1 , not ‖ 𝐴 ‖ 2 . For uniform attention over 𝑁 tokens, ‖ 𝐴 ‖ 2 = 1 / 𝑁 . Variance collapses as context grows. With L1 softmax, output magnitude depends on both sparsity and sequence length. With L2 softmax, it doesn’t. The same holds for gradients. Two identical models, differing only in softmax norm, trained to count occurrences of a target token. Training uses staged seq

Explore this link on the map →

saved by