KernelBench v0.1 | Scaling Intelligence Lab at Stanford University
Since the initial v0 release, the benchmark has been used in a variety of promising ways to study and improve LLMs’ ability to write kernel code. As always, thank you to the community –– including teams from Cognition AI, Meta, METR, Nvidia, Prime Intellect, SakanaAI, and many others whose work we may not yet be aware of –– for working on the benchmark and helping surface both its strengths and limitations. Your feedback directly informed many of the improvements in v0.1 and we’re excited to see what you’ll build with the new version. Hopefully, v0.1 pushes these efforts even further and inspires many more to come! In what follows, we walk through the key changes behind KernelBench v0.1. We start by outlining the most common reward hacking patterns we’ve observed, along with a guideline for checking whether performance claims make actual sense. We then describe the changes we made to improve task quality and testing robustness. So far here are some examples of common reward hacking pat
KernelBench v0.1 | Scaling Intelligence Lab at Stanford University --> Scaling Intelligence Lab Home About --> People --> · Publications · Blogs · Openings · Code --> --> KernelBench v0.1 Natalia Kokoromyti Stanford Sokserey Sun Stanford Anne Ouyang Stanford Azalia Mirhoseini Stanford KernelBench v0.1 is out KernelBench v0.1 TL;DR: KernelBench v0.1 is out What’s new: We provide a guideline that makes it easier to analyze the validity of results and quickly rule out physically impossible performance claims. Support for randomized testing beyond normal distributions.
Explore this link on the map →saved by
related reading
- Kevin-32B: Multi-Turn RL for Writing CUDA Kernels | Cognitioncognition.ai
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- “This Kernel Was Faster Yesterday” — In Pursuit of High-Fidelity GPU Kernel Benchmarkingstandardkernel.com
- How to Land a Frontier Lab Jobvladfeinberg.com
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- Composer2.pdfcursor.com
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- The Reward Hacking Benchmarkkunvarthaman.com
- Strangely, Matrix Multiplications on GPUs Run Faster When Given "Predictable" Data! [short]thonking.ai
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com