AI Benchmarking Is Broken
To celebrate Patronus AI’s $17M Series A, I wanted to share some thoughts on the current state of AI benchmarking. Imagine a standardized test that works like this: In case you’re wondering, the correct reaction is “this is not a very good exam.” Cheating is rampant. The answers can all be memorized. The test isn’t really much of a “test.” The scores are meaningless. This is the current state of AI benchmarking. Ironically, we hold AI benchmarks to a very low standard. We allow practices that would never fly in other serious domains. It’s happening in plain sight, with a nod and a wink every time an AI lab releases a new model that tops the leaderboards. We must do better. Thoughtful analysis of the business and economics of tech The first dirty secret of AI benchmarks is quite literally, dirty: training data contamination. A common refrain among machine learning researchers and engineers — “don’t train on the test set.” In other words, don’t train a model on the same dataset that you
To celebrate Patronus AI’s $17M Series A , I wanted to share some thoughts on the current state of AI benchmarking. Imagine a standardized test that works like this: The questions and answers are freely available on the public internet Cheating is not regulated, it’s a 100% honor system The exam never changes, it’s the same questions every time Scores on the exam seem to be getting better every year In case you’re wondering, the correct reaction is “this is not a very good exam.” Cheating is rampant. The answers can all be memorized. The test isn’t really much of a “test.” The scores are meani
Explore this link on the map →related reading
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- Contra Labs - Powered by Contracontralabs.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- There's An AI For That® — The front page of AItheresanaiforthat.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- AI Benchmark Leaderboards & Model Evals | BenchmarkListbenchmarklist.com
- A statistical approach to model evaluations \ Anthropicanthropic.com
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHubgithub.com
- GitHub - salesforce/AuditNLG: AuditNLG: Auditing Generative AI Language Modeling for Trustworthiness · GitHubgithub.com