Benchmark Scores = General Capability + Claudiness
epochai.substack.com · 1,229 words · saved by 1 readers
Is this because skills generalize very well, or because developers are pushing on all benchmarks at once?
The Gemini 3 release included a massive table showing how the model was state-of-the-art on nineteen diverse benchmarks. Such tables are commonplace by now, but they add up to an odd statistical situation. Benchmarks ostensibly measure different things, but since models tend to improve on many benchmarks at once, the dataset of benchmark scores is dominated by a single “General Capability” dimension. In this post, I’ll describe the statistics of this dataset, look into what’s left when you factor out this dominant dimension (hint: it’s “Claudiness”), and discuss how this relates to an…
related reading
- Benchmark Scores = General Capability + Claudinesssubstack.com
- Epoch Capabilities Indexepoch.ai
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Giovanni D'Antoniogiovannidantonio.com
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com
- gpt-4.pdfcdn.openai.com
- A statistical approach to model evaluations \ Anthropicanthropic.com
- Composer2.pdfcursor.com
- [1911.01547] On the Measure of Intelligencearxiv.org
- CAIS AI Dashboarddashboard.safe.ai
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai