flâneur — a map of the web's best reading

KernelBench v0.1 | Scaling Intelligence Lab at Stanford University

scalingintelligence.stanford.edu · 1,492 words · saved by 1 readers

Since the initial v0 release, the benchmark has been used in a variety of promising ways to study and improve LLMs’ ability to write kernel code. As always, thank you to the community –– including teams from Cognition AI, Meta, METR, Nvidia, Prime Intellect, SakanaAI, and many others whose work we may not yet be aware of –– for working on the benchmark and helping surface both its strengths and limitations. Your feedback directly informed many of the improvements in v0.1 and we’re excited to see what you’ll build with the new version. Hopefully, v0.1 pushes these efforts even further and inspires many more to come! In what follows, we walk through the key changes behind KernelBench v0.1. We start by outlining the most common reward hacking patterns we’ve observed, along with a guideline for checking whether performance claims make actual sense. We then describe the changes we made to improve task quality and testing robustness. So far here are some examples of common reward hacking pat

KernelBench v0.1 | Scaling Intelligence Lab at Stanford University --> Scaling Intelligence Lab Home About --> People --> · Publications · Blogs · Openings · Code --> --> KernelBench v0.1 Natalia Kokoromyti Stanford Sokserey Sun Stanford Anne Ouyang Stanford Azalia Mirhoseini Stanford KernelBench v0.1 is out KernelBench v0.1 TL;DR: KernelBench v0.1 is out What’s new: We provide a guideline that makes it easier to analyze the validity of results and quickly rule out physically impossible performance claims. Support for randomized testing beyond normal distributions.

Explore this link on the map →

saved by

related reading