✳flâneur — a map of the web's best reading
We reverse-engineered Flash Attention 4
modal.com · 4,225 words · saved by 1 readers
Asynchrony, fast approximate exponents, and 10x more efficient softmax.
All posts Back Research September 26, 2025 • 15 minute read We reverse-engineered Flash Attention 4 Charles Frye Developer Advocate Nathan Wang Member of Technical Staff Timothy Feng Member of Technical Staff This blog post made the front page of HackerNews! Discussion here . This blog post was presented to the GPU MODE Discord ! Watch the recording here . One month ago at Hot Chips , Tri Dao presented preliminary results on Flash Attention 4, the latest addition to the Flash Attention series of CUDA kernels . These kernels are used in the attention layers of Transformer neural networks. Along
Explore this link on the map →related reading
- From Online Softmax to FlashAttentioncourses.cs.washington.edu
- Flash Attention from Scratch Part 1: Introlubits.ch
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- GPUs Go Brrr · Hazy Researchhazyresearch.stanford.edu
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- From Pairwise to Higher Order Tensor Operations on GPUs – Springtail Blogspringtail.ai
- [2410.20399] ThunderKittens: Simple, Fast, and Adorable AI Kernelsarxiv.org
- Overleaf Examplearxiv.org
- A User's Guide to FlexAttention in FlashAttention CuTe DSL - Colfax Researchresearch.colfax-intl.com
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev