2307.08691
arxiv.org · 6,149 words · saved by 1 readers
N/A
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning Tri Dao1,2 1 Department of Computer Science, Princeton University 2 Department of Computer Science, Stanford University arXiv:2307.08691v1 [cs.LG] 17 Jul 2023…
related reading
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org
- 2205.14135arxiv.org
- Biao's Bloghebiao064.github.io
- We reverse-engineered Flash Attention 4modal.com
- Flash Attention from Scratch Part 1: Introlubits.ch
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org
- Overleaf Examplearxiv.org
- All the Transformer Math You Need to Know | How To Scale Your Modeljax-ml.github.io
- From Online Softmax to FlashAttentioncourses.cs.washington.edu
- Mamba: The Easy Wayjackcook.com