flâneur — a map of the web's best reading

DeltaNet Explained (Part I) | Songlin Yang

sustcsonglin.github.io · 2,269 words · saved by 4 readers

A gentle and comprehensive introduction to the DeltaNet

DeltaNet Explained (Part I) | Songlin Yang DeltaNet Explained (Part I) A gentle and comprehensive introduction to the DeltaNet This blog post series accompanies our NeurIPS ‘24 paper - Parallelizing Linear Transformers with the Delta Rule over Sequence Length (w/ Bailin Wang , Yu Zhang , Yikang Shen and Yoon Kim ). You can find the implementation here and the presentation slides here . Part I - The Model Part II - The Algorithm Part III - The Neural Architecture Linear attention as RNN Notations: we use CAPITAL BOLD letters to represent matrices, lowercase bold letters to represent vectors, an

Explore this link on the map →

saved by

related reading