The end of benchmarks
Imagine a society of apes trying to play tic tac toe. One day, they meet a very good human player, who also has a computer that has a superintelligent AI on it, trained to play tic tac toe. From their perspective, both the human and the computer perform equally well; they can stalemate each other every time, and they often beat the apes. Even if the superintelligent AI is much better than the human at, stay, Go, the apes could not tell. This is an example of the problem of intelligence saturation. If the hardest task you perform in your society is tic tac toe, you cannot distinguish between two intelligences that both saturate performance on tic tac toe. In our case, we are far from intelligence saturation, but it is certainly possible to imagine that, at some point, we will have models that saturate virtually all tasks that humans perform today, at which point we will no longer be able to tell which model is more intelligent, at least on any tasks we care about. Over the past 6 months
Imagine a society of apes trying to play tic tac toe. One day, they meet a very good human player, who also has a computer that has a superintelligent AI on it, trained to play tic tac toe. From their perspective, both the human and the computer perform equally well; they can stalemate each other every time, and they often beat the apes. Even if the superintelligent AI is much better than the human at, stay, Go, the apes could not tell. This is an example of the problem of intelligence saturation. If the hardest task you perform in your society is tic tac toe, you cannot distinguish between tw
Explore this link on the map →related reading
- [1911.01547] On the Measure of Intelligencearxiv.org
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com
- Infinite midwit - by Adam Mastroianniexperimental-history.com
- After Automation | Everyevery.to
- Artificial General Intelligence Is Already Herenoemamag.com
- My picture of the present in AI — LessWronglesswrong.com
- The Yale Review | Melanie Mitchell: The Dangerous Unknowns at the…yalereview.org
- The bitter lesson of LLM evalsparsed.com
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com