On measuring AI - Nikil Ravi
nikilravi.substack.com · 1,572 words · saved by 1 readers
Thoughts on the state of measurement in AI
Of late, the number of AI model releases per month has been growing very rapidly. Each model’s release is accompanied by the usual benchmark scores that show how well the model performs. Usually, all the charts released by the model provider show their model beating its predecessor and competitors on every reported benchmark. Everything trends up and to the right, this gets reported by various media outlets and reposted by prominent social media accounts, and within a few days, the benchmark scores have either directly or indirectly succeeded in shaping the opinion of most people about the…
saved by
related reading
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- Giovanni D'Antoniogiovannidantonio.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Toward A Public Science of Model Behavior | Transluce AItransluce.org
- Successful language model evals - Jason Weijasonwei.net
- Things I learned at OpenAI - by Karina Nguyen - sémaphoresemaphore.substack.com
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- A statistical approach to model evaluations \ Anthropicanthropic.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- AI’s capabilities may be exaggerated by flawed tests, study saysnbcnews.com
- We Need A ‘Science of Evals’ – Apollo Researchapolloresearch.ai