AMD GPUs go brrr · Hazy Research
hazyresearch.stanford.edu · 4,747 words · saved by 1 readers
multi silicon ai is coming
AMD GPUs go brrr · Hazy Research Nov 11, 2025 · 25 min read AMD GPUs go brrr William Hu , Drew Wadsworth , Simran Arora Team : William Hu, Drew Wadsworth, Sean Siddens, Stanley Winata, Daniel Fu, Ryan Swann, Muhammad Osama, Christopher Ré, Simran Arora Links : Arxiv | Code AI is compute hungry . So we've been asking : How do we build AI from the hardware up? How do we lead AI developers to do what the hardware prefers? AMD GPUs are now offering state-of-the-art speeds and feeds. However, this performance is locked away from AI workflows due to the lack of mature AMD software . We share HipKitt
saved by
related reading
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- GPUs Go Brrr · Hazy Researchhazyresearch.stanford.edu
- Worklog: Optimising GEMM on NVIDIA H100 for cuBLAS-like Performance (WIP) – Hamza's Bloghamzaelshafie.bearblog.dev
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- GitHub - adam-maj/tiny-gpu: A minimal GPU design in Verilog to learn how GPUs work from the ground up · GitHubgithub.com
- [2410.20399] ThunderKittens: Simple, Fast, and Adorable AI Kernelsarxiv.org
- AI Chip Architecturesjacobpeake.com
- Execution Model - SLING user documentationdoc.sling.si
- Outperforming cuBLAS on H100: a Worklogcudaforfun.substack.com
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io
- BrrrVizbrrrviz.com