2305.19370
arxiv.org · 7,400 words · saved by 1 readers
N/A
Blockwise Parallel Transformer for Large Context Models Hao Liu Pieter Abbeel UC Berkeley UC Berkeley hao.liu@cs.berkeley.edu pabbeel@cs.berkeley.edu arXiv:2305.19370v3 [cs.CL] 28 Aug 2023…
related reading
- 2310.01889arxiv.org
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org
- transformer_attention.pdfarxiv.org
- 1706.03762arxiv.org
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- Linear Transformers Are Faster After All – Manifest AImanifestai.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Overleaf Examplearxiv.org
- [2507.04239] Scaling Context Requires Rethinking Attentionarxiv.org
- Transformers from scratch | peterbloem.nlpeterbloem.nl
- [2011.04006] Long Range Arena: A Benchmark for Efficient Transformersarxiv.org