1911.02150
arxiv.org · 4,599 words · saved by 1 readers
N/A
Fast Transformer Decoding: One Write-Head is All You Need Noam Shazeer Google noam@google.com arXiv:1911.02150v1 [cs.NE] 6 Nov 2019 November 7, 2019…
related reading
- transformer_attention.pdfarxiv.org
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- 1706.03762arxiv.org
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- The Annotated Transformernlp.seas.harvard.edu
- The Annotated Transformernlp.seas.harvard.edu
- Multi Query Attention (MQA) and Grouped-Query Attention (GQA)tinkerd.net
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org
- 2305.13245arxiv.org