flâneur

Benchmark Scores = General Capability + Claudiness

epochai.substack.com · 1,229 words · saved by 1 readers

Is this because skills generalize very well, or because developers are pushing on all benchmarks at once?

The Gemini 3 release included a massive table showing how the model was state-of-the-art on nineteen diverse benchmarks. Such tables are commonplace by now, but they add up to an odd statistical situation. Benchmarks ostensibly measure different things, but since models tend to improve on many benchmarks at once, the dataset of benchmark scores is dominated by a single “General Capability” dimension. In this post, I’ll describe the statistics of this dataset, look into what’s left when you factor out this dominant dimension (hint: it’s “Claudiness”), and discuss how this relates to an…

related reading