flâneur — a map of the web's best reading

Parallelizing Linear Transformers with the Delta Rule over Sequence Length

arxiv.org · 27,660 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.

Parallelizing Linear Transformers with the Delta Rule over Sequence Length Songlin Yang ⋄ Bailin Wang ⋄ Yu Zhang † Yikang Shen ‡ Yoon Kim ⋄ ⋄ Massachusetts Institute of Technology † Soochow University ‡ MIT-IBM Watson AI Lab yangsl66@mit.edu Abstract Transformers with linear attention (i.e., linear transformers) and state-space models have recently been suggested as a viable linear-time alternative to transformers with softmax attention. However, these models still underperform transformers especially on tasks that require in-context retrieval. While more expressive variants of linear transfor

Explore this link on the map →

saved by

related reading