KernelBench v0.1 | Scaling Intelligence Lab at Stanford University
Since the initial v0 release, the benchmark has been used in a variety of promising ways to study and improve LLMs’ ability to write kernel code. As always, thank you to the community –– including teams from Cognition AI, Meta, METR, Nvidia, Prime Intellect, SakanaAI, and many others whose work we may not yet be aware of –– for working on the benchmark and helping surface both its strengths and limitations. Your feedback directly informed many of the improvements in v0.1 and we’re excited to see what you’ll build with the new version. Hopefully, v0.1 pushes these efforts even further and inspires many more to come! In what follows, we walk through the key changes behind KernelBench v0.1. We start by outlining the most common reward hacking patterns we’ve observed, along with a guideline for checking whether performance claims make actual sense. We then describe the changes we made to improve task quality and testing robustness. So far here are some examples of common reward hacking pat
KernelBench v0.1 | Scaling Intelligence Lab at Stanford University --> Scaling Intelligence Lab Home About --> People --> · Publications · Blogs · Openings · Code --> --> KernelBench v0.1 Natalia Kokoromyti Stanford Sokserey Sun Stanford Anne Ouyang Stanford Azalia Mirhoseini Stanford KernelBench v0.1 is out KernelBench v0.1 TL;DR: KernelBench v0.1 is out What’s new: We provide a guideline that makes it easier to analyze the validity of results and quickly rule out physically impossible performance claims. Support for randomized testing beyond normal distributions.
saved by
related reading
- KernelBench: Can LLMs Write GPU Kernels?scalingintelligence.stanford.edu
- Kevin-32B: Multi-Turn RL for Writing CUDA Kernels | Cognitioncognition.ai
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- How to Land a Frontier Lab Jobvladfeinberg.com
- “This Kernel Was Faster Yesterday” — In Pursuit of High-Fidelity GPU Kernel Benchmarkingstandardkernel.com
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- Every Benchmark is Brokenjonathanpgabor.substack.com
- GitHub - wafer-ai/gpu-perf-engineering-resources: A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.github.com
- Composer2.pdfcursor.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com