✳flâneur — a map of the web's best reading
GPUs Go Brrr · Hazy Research
hazyresearch.stanford.edu · 4,846 words · saved by 1 readers
how make gpu fast?
GPUs Go Brrr · Hazy Research May 12, 2024 · 26 min read GPUs Go Brrr Benjamin Spector , Aaryan Singhal , Simran Arora , Chris Re AI uses an awful lot of compute . In the last few years we’ve focused a great deal of our work on making AI use less compute (e.g. Based , Monarch Mixer , H3 , Hyena , S4 , among others) and run more efficiently on the compute that we have (e.g. FlashAttention , FlashAttention-2 , FlashFFTConv ). Lately, reflecting on these questions has prompted us to take a step back, and ask two questions: What does the hardware actually want? And how can we give that to it? This
Explore this link on the map →related reading
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- [2410.20399] ThunderKittens: Simple, Fast, and Adorable AI Kernelsarxiv.org
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Worklog: Optimising GEMM on NVIDIA H100 for cuBLAS-like Performance (WIP) – Hamza's Bloghamzaelshafie.bearblog.dev
- Outperforming cuBLAS on H100: a Worklogcudaforfun.substack.com
- Outperforming cuBLAS on H100: a Worklogcudaforfun.substack.com
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- AMD GPUs go brrr · Hazy Researchhazyresearch.stanford.edu
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io
- Flash Attention from Scratch Part 1: Introlubits.ch
- We reverse-engineered Flash Attention 4modal.com
- CVPR2023_eff_tutorial_molchanov.pdfnvlabs.github.io