Kimi Linear: An Expressive, Efficient Attention Architecture | alphaXiv
Select any part of the paper to ask specific questions Type @ to reference other papers and expand the discussion • See how others cite this work • Literature reviews • Community context Try asking "What's the intuition behind section 3.2?"
Submitted 01 Nov 2025 Abstract We introduce Kimi Linear, a hybrid linear attention architecture that, for the first time, outperforms full attention under fair comparisons across various scenarios -- including short-context, long-context, and reinforcement learning (RL) scaling regimes. At its core lies Kimi Delta Attention (KDA), an expressive linear attention module that extends Gated DeltaNet with a finer-grained gating mechanism, enabling more effective use of limited finite-state RNN memory. Our bespoke chunkwise algorithm achieves high hardware efficiency through a specialized…
saved by
related reading
- Kimi Linear: An Expressive, Efficient Attention Architecturearxiv.org
- [2510.26692] Kimi Linear: An Expressive, Efficient Attention Architecturearxiv.org
- ali (@waterloo_intern) on Xx.com
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- Linear Attention Fundamentals | Hailey Schoelkopfhaileyschoelkopf.github.io
- DeltaNet Explained (Part I) | Songlin Yangsustcsonglin.github.io
- [2607.07953] Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routingarxiv.org
- Log-Linear Attentionarxiv.org
- Linear Transformers Are Faster After All – Manifest AImanifestai.com
- 2502.11089arxiv.org
- Subquadratic — How SSA Makes Long Context Practicalsubq.ai
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io