flâneur — a map of the web's best reading

Reverse-Engineering cuBLAS

accu.org · 2,734 words · saved by 1 readers

It’s possible to achieve cuBLAS performance with tensor cores by mimicking SASS instructions. Fabian Schuetze guides us through the process. A5000 (GPU): A GPU produced by Nvidia. The A5000 is based on the Ampere microarchitecture. The article uses specialized instructions introduced with Ampere. The subsequent microarchitecture (Hopper) introduced new instructions to attain maximum performance on these types of GPUs. BLAS (and GEMM): GEMM stands for General Matrix Multiplication. Refers to a group of operations (called Level 3) of the Basic Linear Algebra Subprograms (BLAS) too. A standardized interface to BLAS will become part of C++ 26 (std::linalg) as proposed by P1673. cuBLAS: Nvidia’s variant of the BLAS library. It contains highly optimized and specialized code for all GPU variants and matrix sizes. Its source code is not publicly accessible. CUDA: An extension of the C language to write programs for Nvidia GPUs. CUDA affords programmers the ability to control the L1 cache of su

Reverse-Engineering cuBLAS Reverse-Engineering cuBLAS Reverse-Engineering cuBLAS By Fabian Schuetze Overload, 32(181):9-13, June 2024 It’s possible to achieve cuBLAS performance with tensor cores by mimicking SASS instructions. Fabian Schuetze guides us through the process. Glossary A5000 (GPU) : A GPU produced by Nvidia. The A5000 is based on the Ampere microarchitecture. The article uses specialized instructions introduced with Ampere. The subsequent microarchitecture (Hopper) introduced new instructions to attain maximum performance on these types of GPUs. BLAS (and GEMM ): GEMM stands for

Explore this link on the map →

saved by

related reading