✳flâneur — a map of the web's best reading
AMD GPUs go brrr · Hazy Research
hazyresearch.stanford.edu · 4,747 words · saved by 1 readers
multi silicon ai is coming
AMD GPUs go brrr · Hazy Research Nov 11, 2025 · 25 min read AMD GPUs go brrr William Hu , Drew Wadsworth , Simran Arora Team : William Hu, Drew Wadsworth, Sean Siddens, Stanley Winata, Daniel Fu, Ryan Swann, Muhammad Osama, Christopher Ré, Simran Arora Links : Arxiv | Code AI is compute hungry . So we've been asking : How do we build AI from the hardware up? How do we lead AI developers to do what the hardware prefers? AMD GPUs are now offering state-of-the-art speeds and feeds. However, this performance is locked away from AI workflows due to the lack of mature AMD software . We share HipKitt
Explore this link on the map →saved by
related reading
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- GPUs Go Brrr · Hazy Researchhazyresearch.stanford.edu
- Worklog: Optimising GEMM on NVIDIA H100 for cuBLAS-like Performance (WIP) – Hamza's Bloghamzaelshafie.bearblog.dev
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- [2410.20399] ThunderKittens: Simple, Fast, and Adorable AI Kernelsarxiv.org
- Execution Model - SLING user documentationdoc.sling.si
- Outperforming cuBLAS on H100: a Worklogcudaforfun.substack.com
- BrrrVizbrrrviz.com
- Outperforming cuBLAS on H100: a Worklogcudaforfun.substack.com
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io
- CUDA C++ Programming Guide (Legacy) — CUDA C++ Programming Guidedocs.nvidia.com