Your Evals Will Break and You Won't See It Coming - Lun Wang
We're good at evaluating the models we have. We're much worse at evaluating the models we're about to build — especially if they cross into a new capability regime. Most benchmarks, safety evals, and red-teaming protocols implicitly assume the next model is a stronger version of the current one. If it's a different kind of thing, our entire evaluation infrastructure breaks silently. I think this is the most important unsolved problem in how we understand LLMs. And I think the answer is that eval — not training, not architecture, not data — is the bottleneck for the next capability jump. Let me explain why. Wei et al. (2022) documented what they called "emergent abilities" — few-shot prompted task performance, chain-of-thought reasoning gains, instruction following — capabilities that appeared only at larger scales. Grokking (Power et al., 2022) shows a related but distinct phenomenon: networks that suddenly generalize long after memorizing their training data, a dynamic transition over
We're good at evaluating the models we have. We're much worse at evaluating the models we're about to build - especially if they cross into a new capability regime. Most benchmarks, safety evals, and red-teaming protocols implicitly assume the next model is a stronger version of the current one. If it's a different kind of thing, our entire evaluation infrastructure breaks silently. I think this is the most important unsolved problem in how we understand LLMs. And I think the answer is that eval - not training, not architecture, not data - is the bottleneck for the next capability jump. Let me
Explore this link on the map →saved by
related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- The bitter lesson of LLM evalsparsed.com
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com
- Things I learned at OpenAI - by Karina Nguyen - sémaphoresemaphore.substack.com
- Model evals for dangerous capabilities — LessWronglesswrong.com
- Predicting LLM Safety Before Release by Simulating Deploymentcdn.openai.com
- AI in 2025: gestalt — LessWronglesswrong.com
- A statistical approach to model evaluations \ Anthropicanthropic.com
- LLM evaluation: a beginner's guideevidentlyai.com
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- A starter guide for evals — AI Alignment Forumalignmentforum.org
- Challenges in evaluating AI systems \ Anthropicanthropic.com