Reverse-Engineering cuBLAS
It’s possible to achieve cuBLAS performance with tensor cores by mimicking SASS instructions. Fabian Schuetze guides us through the process. A5000 (GPU): A GPU produced by Nvidia. The A5000 is based on the Ampere microarchitecture. The article uses specialized instructions introduced with Ampere. The subsequent microarchitecture (Hopper) introduced new instructions to attain maximum performance on these types of GPUs. BLAS (and GEMM): GEMM stands for General Matrix Multiplication. Refers to a group of operations (called Level 3) of the Basic Linear Algebra Subprograms (BLAS) too. A standardized interface to BLAS will become part of C++ 26 (std::linalg) as proposed by P1673. cuBLAS: Nvidia’s variant of the BLAS library. It contains highly optimized and specialized code for all GPU variants and matrix sizes. Its source code is not publicly accessible. CUDA: An extension of the C language to write programs for Nvidia GPUs. CUDA affords programmers the ability to control the L1 cache of su
Reverse-Engineering cuBLAS Reverse-Engineering cuBLAS Reverse-Engineering cuBLAS By Fabian Schuetze Overload, 32(181):9-13, June 2024 It’s possible to achieve cuBLAS performance with tensor cores by mimicking SASS instructions. Fabian Schuetze guides us through the process. Glossary A5000 (GPU) : A GPU produced by Nvidia. The A5000 is based on the Ampere microarchitecture. The article uses specialized instructions introduced with Ampere. The subsequent microarchitecture (Hopper) introduced new instructions to attain maximum performance on these types of GPUs. BLAS (and GEMM ): GEMM stands for
Explore this link on the map →saved by
related reading
- Outperforming cuBLAS on H100: a Worklogcudaforfun.substack.com
- CUTLASS: Fast Linear Algebra in CUDA C++ | NVIDIA Technical Blogdeveloper.nvidia.com
- GitHub - wangzyon/NVIDIA_SGEMM_PRACTICE: Step-by-step optimization of CUDA SGEMM · GitHubgithub.com
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Outperforming cuBLAS on H100: a Worklogcudaforfun.substack.com
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- GPU Performance Background User's Guide - NVIDIA Docsdocs.nvidia.com
- Worklog: Optimising GEMM on NVIDIA H100 for cuBLAS-like Performance (WIP) – Hamza's Bloghamzaelshafie.bearblog.dev
- Mini Project: GPU Accelerated Matrix Multiplication (almost) like cuBLAS0mean1sigma.com
- Implementing a fast Tensor Core matmul on the Ada Architecture | spatters.caspatters.ca
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- Learning CUDA by optimizing matrix-vector multiplication (SGEMV) for cuBLAS-like performance - A worklog – Maharshi's blogmaharshi.bearblog.dev