flash-attention/csrc/flash_attn/src/flash_fwd_kernel.h at main · Dao-AILab/flash-attention
github.com · 8,835 words · saved by 1 readers
Fast and memory-efficient exact attention. Contribute to Dao-AILab/flash-attention development by creating an account on GitHub.
saved by
related reading
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org
- We reverse-engineered Flash Attention 4modal.com
- Biao's Bloghebiao064.github.io
- Flash Attention from Scratch Part 1: Introlubits.ch
- 2307.08691arxiv.org
- From Online Softmax to FlashAttentioncourses.cs.washington.edu
- Overleaf Examplearxiv.org
- Attention-Residuals/Attention_Residuals.pdf at master · MoonshotAI/Attention-Residuals · GitHubgithub.com
- A User's Guide to FlexAttention in FlashAttention CuTe DSL - Colfax Researchresearch.colfax-intl.com
- 2205.14135arxiv.org
- 2502.11089arxiv.org
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org