2310.01889
arxiv.org · 7,754 words · saved by 1 readers
N/A
Ring Attention with Blockwise Transformers for Near-Infinite Context Hao Liu, Matei Zaharia, Pieter Abbeel UC Berkeley hao.liu@cs.berkeley.edu arXiv:2310.01889v4 [cs.CL] 27 Nov 2023 Abstract…
related reading
- 2305.19370arxiv.org
- [2507.04239] Scaling Context Requires Rethinking Attentionarxiv.org
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org
- Overleaf Examplearxiv.org
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Linear Transformers Are Faster After All – Manifest AImanifestai.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- Mediumblog.gopenai.com
- LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelismarxiv.org
- 2205.14135arxiv.org
- [2011.04006] Long Range Arena: A Benchmark for Efficient Transformersarxiv.org
- Star Attention: Efficient LLM Inference over Long Sequencesarxiv.org