✳flâneur — a map of the web's best reading
CVPR2023_eff_tutorial_molchanov.pdf
nvlabs.github.io · 3,029 words · saved by 1 readers
N/A
# link_2ao4p8wm4ip.pdf ## Metadata - PDFFormatVersion=1.4 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Producer=macOS Version 13.4 (Build 22F66) Quartz PDFContext - CreationDate=D:20230619223826Z00'00' - ModDate=D:20230619223826Z00'00' ## Contents ### Page 1 1NN Performance optimization:How to achieve more with less cost Software perspectiveDisclaimer: Results, numbers and performance are reported from the research perspective. For the exact performance please contact NVIDIA product managers. Pavlo Molchanov, p
Explore this link on the map →saved by
related reading
- GPU Performance Background User's Guide - NVIDIA Docsdocs.nvidia.com
- Worklog: Optimising GEMM on NVIDIA H100 for cuBLAS-like Performance (WIP) – Hamza's Bloghamzaelshafie.bearblog.dev
- NVIDIA Tensor Core Evolution: From Volta To Blackwellnewsletter.semianalysis.com
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- How to Think About TPUs | How To Scale Your Modeljax-ml.github.io
- All About Rooflines | How To Scale Your Modeljax-ml.github.io
- Linear/Fully-Connected Layers User's Guide - NVIDIA Docsdocs.nvidia.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Making Deep Learning go Brrrr From First Principleshorace.io
- How To Scale Your Modeljax-ml.github.io