Epoch Capabilities Index | Epoch AI
epoch.ai · 334 words · saved by 1 readers
The Epoch Capabilities Index combines many benchmarks into a single capability scale for comparing models over time.
The general ECI is a composite metric which uses scores from over 50 distinct benchmarks to generate a single, general capability scale. At a high level, ECI stitches together its component benchmarks, determining their relative difficulty by making comparisons wherever models are evaluated on multiple benchmarks. Individual models obtain higher ECI scores if they perform better on harder benchmarks. We give an overview of our methodology here; further technical details are available in our paper, A Rosetta Stone for AI Benchmarks, which was funded by Google DeepMind, and written in…
saved by
related reading
- Frontier AI Cybersecurity Observatorycybergym.io
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- CAIS AI Dashboarddashboard.safe.ai
- My picture of the present in AI — LessWronglesswrong.com
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Benchmark Scores = General Capability + Claudinesssubstack.com
- Benchmark Scores = General Capability + Claudinessepochai.substack.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- AI progress is about to speed up | Epoch AIepoch.ai
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- Measuring AI Ability to Complete Long Tasks - METRmetr.org