flâneur — a map of the web's best reading

AI Benchmarking Is Broken

whoisnnamdi.com · 163 words · saved by 1 readers

To celebrate Patronus AI’s $17M Series A, I wanted to share some thoughts on the current state of AI benchmarking. Imagine a standardized test that works like this: In case you’re wondering, the correct reaction is “this is not a very good exam.” Cheating is rampant. The answers can all be memorized. The test isn’t really much of a “test.” The scores are meaningless. This is the current state of AI benchmarking. Ironically, we hold AI benchmarks to a very low standard. We allow practices that would never fly in other serious domains. It’s happening in plain sight, with a nod and a wink every time an AI lab releases a new model that tops the leaderboards. We must do better. Thoughtful analysis of the business and economics of tech The first dirty secret of AI benchmarks is quite literally, dirty: training data contamination. A common refrain among machine learning researchers and engineers — “don’t train on the test set.” In other words, don’t train a model on the same dataset that you

To celebrate Patronus AI’s $17M Series A , I wanted to share some thoughts on the current state of AI benchmarking. Imagine a standardized test that works like this: The questions and answers are freely available on the public internet Cheating is not regulated, it’s a 100% honor system The exam never changes, it’s the same questions every time Scores on the exam seem to be getting better every year In case you’re wondering, the correct reaction is “this is not a very good exam.” Cheating is rampant. The answers can all be memorized. The test isn’t really much of a “test.” The scores are meani

Explore this link on the map →

related reading