A consumption basket approach to measuring AI progress - Marginal REVOLUTION
Many AI evaluations go out of their way to find hard problems. That makes sense because you can track progress over time, and furthermore many of the world’s important problems are hard problems, such as building out advances in the biosciences. One common approach, for instance, is to track the performance of current AI models on say International Math Olympiad problems. I am all for those efforts, and I do not wish to cut back on them. Still, they introduce biases in our estimates of progress. Many of those measures show that the AIs still are not solving most of the core problems, and sometimes they are not coming close. In contrast, actual human users typically deploy AIs to help them with relatively easy problems. They use AIs for (standard) legal advice, to help with the homework, to plot travel plans, to help modify a recipe, as a therapist or advisor, and so on. You could say that is the actual consumption basket for LLM use, circa 2025. It would be interesting to chart the
Many AI evaluations go out of their way to find hard problems. That makes sense because you can track progress over time, and furthermore many of the world’s important problems are hard problems, such as building out advances in the biosciences. One common approach, for instance, is to track the performance of current AI models on say International Math Olympiad problems. I am all for those efforts, and I do not wish to cut back on them. Still, they introduce biases in our estimates of progress. Many of those measures show that the AIs still are not solving most of the core problems, and
Explore this link on the map →related reading
- Economics and Transformative AI | Tom Cunningham – Tom Cunninghamtecunningham.github.io
- AI as Normal Technology | Knight First Amendment Instituteknightcolumbia.org
- AI in 2025: gestalt — LessWronglesswrong.com
- My picture of the present in AI — LessWronglesswrong.com
- AI progress is about to speed up | Epoch AIepoch.ai
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- How fast is AI improving? - AI Digesttheaidigest.org
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- AI #100: Meet the New Boss | Don't Worry About the Vasethezvi.wordpress.com
- State of AI 2025: 100T Token LLM Usage Study | OpenRouteropenrouter.ai
- The least understood driver of AI progress | Epoch AIepoch.ai