Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
arxiv.org · 6,276 words · saved by 1 readers
N/A
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention Angelos Katharopoulos 1 2 Apoorv Vyas 1 2 Nikolaos Pappas 3 François Fleuret 2 4 * Abstract by the global receptive field of self-attention, which pro- arXiv:2006.16236v3 [cs.LG] 31 Aug 2020 Transformers achieve remarkable performance in…
related reading
- transformer_attention.pdfarxiv.org
- Linear Transformers Are Faster After All – Manifest AImanifestai.com
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- 1706.03762arxiv.org
- Transformers from Scratche2eml.school
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- Linear Attention Is All You Need | Towards Data Sciencetowardsdatascience.com
- Transformers from scratch | peterbloem.nlpeterbloem.nl
- The Annotated Transformernlp.seas.harvard.edu
- [2011.04006] Long Range Arena: A Benchmark for Efficient Transformersarxiv.org
- Log-Linear Attentionarxiv.org
- [1807.03819] Universal Transformersarxiv.org