wafer-ai/gpu-perf-engineering-resources: A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference. ·
github.com · 1,742 words · saved by 2 readers
A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference. - wafer-ai/gpu-perf-engineering-resources
A curated resource list for learning GPU performance engineering and production inference. The list is ordered from a single inference request to a single GPU, optimized kernels, inference engines, and distributed systems. Read Start here first. After that, use it as a reference. The core list uses original papers, official specifications and documentation, creator repositories, and direct implementation work. If you work on these problems, Wafer is hiring. Contents Start here: the minimum mental model 1. GPU fundamentals Programming model Compilation and machine code 2. Kernel…
saved by
related reading
- Modern GPU Programming For MLSys — Modern GPU Programming For MLSysmlc.ai
- KernelBench: Can LLMs Write GPU Kernels?scalingintelligence.stanford.edu
- Machine Learning System Resources | std::bodun::blogbodunhu.com
- AI Chip Architecturesjacobpeake.com
- How To Scale Your Modeljax-ml.github.io
- GitHub - adam-maj/tiny-gpu: A minimal GPU design in Verilog to learn how GPUs work from the ground up · GitHubgithub.com
- Together AI | The AI Native Cloudtogether.ai
- GitHub - gpu-mode/resource-stream: GPU programming related news and material linksgithub.com
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- Real-time LLM Inference on Standard Datacenter GPUs (3,000 tokens/s per request)blog.kog.ai
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- Modern GPU Programming For MLSys — Modern GPU Programming For MLSysmlc.ai