Devising ML Metrics | CAIS
Metrics drive the ML field. As such, if we want to influence the field or popularize new subfields, we must define the metrics that correlate with progress on the problems we care about. Formalizing these metrics into benchmarks will be crucial to capturing the attention of researchers and driving progress. Building good benchmarks is difficult, in large part because benchmarks exhibit many of the properties that produce power law outcomes. First, implicit in every benchmark design are a large number of multiplicative processes. If even one facet (e.g. ease of use, cost to evaluate, connection of the benchmark with a real problem, tractability, difficulty to game, feasibility of developing new methods to improve the state-of-the-art, etc.) of the benchmark is deficient, it may entirely prevent the benchmark from having impact. Second, benchmarks face strong preferential attachment dynamics: the most used benchmarks are the most likely to be used further. Finally, benchmarks are inheren
Devising ML Metrics | CAIS About About AI risk Resources Resources Contact Careers Donate Our Work Resources AI Risk Contact Careers Donate Careers Donate Devising ML Metrics BLOG AI Risks February 15, 2024 8 min read View as PDF Author: Dan Hendrycks Thomas Woodside Related Posts: A Significant Increase in Digital Labor Automation Submit Your Toughest Questions for Humanity's Last Exam Metrics drive the ML field. As such, if we want to influence the field or popularize new subfields, we must define the metrics that correlate with progress on the problems we care about. Formalizing these metri
Explore this link on the map →related reading
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- Preface - The Emerging Science of Machine Learning Benchmarksmlbenchmarks.org
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- A Bird's Eye View of the ML Field [Pragmatic AI Safety #2] — AI Alignment Forumalignmentforum.org
- Demystifying evals for AI agents \ Anthropicanthropic.com
- [1911.01547] On the Measure of Intelligencearxiv.org
- We Need Better Benchmarks for Machine Learning in Drug Discoverypracticalcheminformatics.blogspot.com
- AI’s capabilities may be exaggerated by flawed tests, study saysnbcnews.com
- Mapping global dynamics of benchmark creation and saturation in artificial intelligence | Nature Communicationsnature.com
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHubgithub.com
- A 5-Minute Guide to UX Benchmarking - by UXPinuxpin.com
- RIP Classic Reasoning Benchmarks. What’s Next?epochai.substack.com