flâneur — a map of the web's best reading

Your Evals Will Break and You Won't See It Coming - Lun Wang

wanglun1996.github.io · 1,258 words · saved by 7 readers

We're good at evaluating the models we have. We're much worse at evaluating the models we're about to build — especially if they cross into a new capability regime. Most benchmarks, safety evals, and red-teaming protocols implicitly assume the next model is a stronger version of the current one. If it's a different kind of thing, our entire evaluation infrastructure breaks silently. I think this is the most important unsolved problem in how we understand LLMs. And I think the answer is that eval — not training, not architecture, not data — is the bottleneck for the next capability jump. Let me explain why. Wei et al. (2022) documented what they called "emergent abilities" — few-shot prompted task performance, chain-of-thought reasoning gains, instruction following — capabilities that appeared only at larger scales. Grokking (Power et al., 2022) shows a related but distinct phenomenon: networks that suddenly generalize long after memorizing their training data, a dynamic transition over

We're good at evaluating the models we have. We're much worse at evaluating the models we're about to build - especially if they cross into a new capability regime. Most benchmarks, safety evals, and red-teaming protocols implicitly assume the next model is a stronger version of the current one. If it's a different kind of thing, our entire evaluation infrastructure breaks silently. I think this is the most important unsolved problem in how we understand LLMs. And I think the answer is that eval - not training, not architecture, not data - is the bottleneck for the next capability jump. Let me

Explore this link on the map →

saved by

related reading