twill.pdf
compilers.stanford.edu · 6,726 words · saved by 1 readers
N/A
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs Rupanshu Soi∗† Rohan Yadav∗† Fredrik Kjolstad Alex Aiken† Stanford University Stanford University Stanford University Stanford University Maryam Mehri Dehnavi Michael Garland Michael Bauer…
saved by
related reading
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- Execution Model - SLING user documentationdoc.sling.si
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- GitHub - adam-maj/tiny-gpu: A minimal GPU design in Verilog to learn how GPUs work from the ground up · GitHubgithub.com
- TPU Deep Divehenryhmko.github.io
- We reverse-engineered Flash Attention 4modal.com
- AMD GPUs go brrr · Hazy Researchhazyresearch.stanford.edu
- Tiny TPUtinytpu.com
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io
- Touching the Elephant - TPUs | Consider the Bulldogconsiderthebulldog.com
- Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programsarxiv.org
- Modern GPU Programming For MLSys — Modern GPU Programming For MLSysmlc.ai