DeltaNet Explained (Part I) | Songlin Yang
sustcsonglin.github.io · 2,269 words · saved by 4 readers
A gentle and comprehensive introduction to the DeltaNet
DeltaNet Explained (Part I) | Songlin Yang DeltaNet Explained (Part I) A gentle and comprehensive introduction to the DeltaNet This blog post series accompanies our NeurIPS ‘24 paper - Parallelizing Linear Transformers with the Delta Rule over Sequence Length (w/ Bailin Wang , Yu Zhang , Yikang Shen and Yoon Kim ). You can find the implementation here and the presentation slides here . Part I - The Model Part II - The Algorithm Part III - The Neural Architecture Linear attention as RNN Notations: we use CAPITAL BOLD letters to represent matrices, lowercase bold letters to represent vectors, an
saved by
related reading
- Linear Attention Fundamentals | Hailey Schoelkopfhaileyschoelkopf.github.io
- [2501.00663] Titans: Learning to Memorize at Test Timearxiv.org
- Test-Time Training with KV Binding Is Secretly Linear Attentionresearch.nvidia.com
- [2412.06464] Gated Delta Networks: Improving Mamba2 with Delta Rulearxiv.org
- Parallelizing Linear Transformers with the Delta Rule over Sequence Lengtharxiv.org
- [2607.07953] Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routingarxiv.org
- DeltaNet Explained (Part II) | Songlin Yangsustcsonglin.github.io
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- ali (@waterloo_intern) on Xx.com
- Log-Linear Attentionarxiv.org
- Kimi Linear: An Expressive, Efficient Attention Architecturearxiv.org
- Learning to (Learn at Test Time): RNNs with Expressive Hidden Statesarxiv.org