✳flâneur — a map of the web's best reading
DeltaNet Explained (Part I) | Songlin Yang
sustcsonglin.github.io · 2,269 words · saved by 4 readers
A gentle and comprehensive introduction to the DeltaNet
DeltaNet Explained (Part I) | Songlin Yang DeltaNet Explained (Part I) A gentle and comprehensive introduction to the DeltaNet This blog post series accompanies our NeurIPS ‘24 paper - Parallelizing Linear Transformers with the Delta Rule over Sequence Length (w/ Bailin Wang , Yu Zhang , Yikang Shen and Yoon Kim ). You can find the implementation here and the presentation slides here . Part I - The Model Part II - The Algorithm Part III - The Neural Architecture Linear attention as RNN Notations: we use CAPITAL BOLD letters to represent matrices, lowercase bold letters to represent vectors, an
Explore this link on the map →saved by
related reading
- Linear Attention Fundamentals | Hailey Schoelkopfhaileyschoelkopf.github.io
- [2501.00663] Titans: Learning to Memorize at Test Timearxiv.org
- Parallelizing Linear Transformers with the Delta Rule over Sequence Lengtharxiv.org
- DeltaNet Explained (Part II) | Songlin Yangsustcsonglin.github.io
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- A short note on some aspects of long context attention | nor's blognor-blog.pages.dev
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- NL.pdfabehrouz.github.io
- [2501.12352] Test-time regression: a unifying framework for designing sequence models with associative memoryar5iv.labs.arxiv.org
- Linear Attention Is All You Need | Towards Data Sciencetowardsdatascience.com
- Subquadratic — How SSA Makes Long Context Practicalsubq.ai
- Post-Transformers - Hyena Hierarchy - by Alex Mackenziewhynowtech.substack.com