[2507.04239] Scaling Context Requires Rethinking Attention
Abstract:We argue that neither transformers nor sub-quadratic architectures are well suited to training at long sequence lengths: the cost of processing the context is too expensive in the former, too inexpensive in the latter. Approaches such as sliding window attention which reduce the cost-per-token of a transformer impair in-context learning, and so are also unsuitable. To address these limitations, we introduce power attention, an architectural layer for linear-cost sequence modeling whose state size can be adjusted independently of parameters, unlocking the advantages of linear attention on practical domains. We develop and open-source a set of GPU kernels for efficient power attention, identifying a novel pattern of operation fusion to avoid memory and bandwidth bottlenecks. Our experiments on the in-context learning of power attention shows that these models dominate both exponential attention and linear attention at long-context training.
Abstract:We argue that neither transformers nor sub-quadratic architectures are well suited to training at long sequence lengths: the cost of processing the context is too expensive in the former, too inexpensive in the latter. Approaches such as sliding window attention which reduce the cost-per-token of a transformer impair in-context learning, and so are also unsuitable. To address these limitations, we introduce power attention, an architectural layer for linear-cost sequence modeling whose state size can be adjusted independently of parameters, unlocking the advantages of linear attention
Explore this link on the map →saved by
related reading
- A short note on some aspects of long context attention | nor's blognor-blog.pages.dev
- Subquadratic — How SSA Makes Long Context Practicalsubq.ai
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Linear Transformers Are Faster After All – Manifest AImanifestai.com
- Extending Context is Hard | kaiokendevkaiokendev.github.io
- transformer_attention.pdfarxiv.org
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- Linear Attention Fundamentals | Hailey Schoelkopfhaileyschoelkopf.github.io
- Mediumblog.gopenai.com
- Overleaf Examplearxiv.org
- From Deep to Long Learning? · Hazy Researchhazyresearch.stanford.edu
- Linear Attention Is All You Need | Towards Data Sciencetowardsdatascience.com