KernelBench: Can LLMs Write GPU Kernels? | Scaling Intelligence Lab at Stanford University
A benchmark designed to evaluate the ability of LLMs to generate efficient GPU kernels for optimizing neural network performance We introduce KernelBench, a benchmark designed to evaluate the ability of large language models (LLMs) to generate efficient GPU kernels for optimizing neural network performance. With 250 well-defined neural network tasks spanning foundational operators, simple fusion patterns, and full ML architectures, the benchmark tasks LLMs to replace PyTorch implementations with custom kernels that are correct and performant. KernelBench highlights the potential for agentic optimization for computer systems with dense feedback signal, where systems iteratively refine kernel designs using profiling tools and tight feedback loops to achieve near-peak hardware utilization. As models scale, well-optimized kernels have far-reaching implications, from reducing the massive energy demands of AI systems to enabling fair and efficient comparisons of novel architectures. By provi
A benchmark designed to evaluate the ability of LLMs to generate efficient GPU kernels for optimizing neural network performance TL;DR We introduce KernelBench, a benchmark designed to evaluate the ability of large language models (LLMs) to generate efficient GPU kernels for optimizing neural network performance. With 250 well-defined neural network tasks spanning foundational operators, simple fusion patterns, and full ML architectures, the benchmark tasks LLMs to replace PyTorch implementations with custom kernels that are correct and performant. KernelBench highlights the potential for…
saved by
related reading
- GitHub - wafer-ai/gpu-perf-engineering-resources: A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.github.com
- Modern GPU Programming For MLSys — Modern GPU Programming For MLSysmlc.ai
- How to Land a Frontier Lab Jobvladfeinberg.com
- KernelBench v0.1 | Scaling Intelligence Lab at Stanford Universityscalingintelligence.stanford.edu
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- About Us - Colfax Researchresearch.colfax-intl.com
- Kevin-32B: Multi-Turn RL for Writing CUDA Kernels | Cognitioncognition.ai
- PostTrainBenchposttrainbench.com
- How To Scale Your Modeljax-ml.github.io
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- Composer2.pdfcursor.com
- Machine Learning System Resources | std::bodun::blogbodunhu.com