✳flâneur — a map of the web's best reading
flash-attention/csrc/flash_attn/src/flash_fwd_kernel.h at main · Dao-AILab/flash-attention
github.com · 8,835 words · saved by 1 readers
Fast and memory-efficient exact attention. Contribute to Dao-AILab/flash-attention development by creating an account on GitHub.
Explore this link on the map →saved by
related reading
- Flash Attention from Scratch Part 1: Introlubits.ch
- From Online Softmax to FlashAttentioncourses.cs.washington.edu
- Attention-Residuals/Attention_Residuals.pdf at master · MoonshotAI/Attention-Residuals · GitHubgithub.com
- Overleaf Examplearxiv.org
- A User's Guide to FlexAttention in FlashAttention CuTe DSL - Colfax Researchresearch.colfax-intl.com
- 2502.11089arxiv.org
- Linear Attention Fundamentals | Hailey Schoelkopfhaileyschoelkopf.github.io
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org
- Mamba: The Easy Wayjackcook.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- Future leakage in block-quantized attention | MatXmatx.com