Devising ML Metrics | CAIS
Metrics drive the ML field. As such, if we want to influence the field or popularize new subfields, we must define the metrics that correlate with progress on the problems we care about. Formalizing these metrics into benchmarks will be crucial to capturing the attention of researchers and driving progress. Building good benchmarks is difficult, in large part because benchmarks exhibit many of the properties that produce power law outcomes. First, implicit in every benchmark design are a large number of multiplicative processes. If even one facet (e.g. ease of use, cost to evaluate, connection of the benchmark with a real problem, tractability, difficulty to game, feasibility of developing new methods to improve the state-of-the-art, etc.) of the benchmark is deficient, it may entirely prevent the benchmark from having impact. Second, benchmarks face strong preferential attachment dynamics: the most used benchmarks are the most likely to be used further. Finally, benchmarks are inheren
Devising ML Metrics | CAIS About About AI risk Resources Resources Contact Careers Donate Our Work Resources AI Risk Contact Careers Donate Careers Donate Devising ML Metrics BLOG AI Risks February 15, 2024 8 min read View as PDF Author: Dan Hendrycks Thomas Woodside Related Posts: A Significant Increase in Digital Labor Automation Submit Your Toughest Questions for Humanity's Last Exam Metrics drive the ML field. As such, if we want to influence the field or popularize new subfields, we must define the metrics that correlate with progress on the problems we care about. Formalizing these metri
related reading
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- Giovanni D'Antoniogiovannidantonio.com
- Successful language model evals - Jason Weijasonwei.net
- Preface - The Emerging Science of Machine Learning Benchmarksmlbenchmarks.org
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- f316275b44ee2de533102913828a8107-Paper-Datasets_and_Benchmarks_Track.pdfproceedings.neurips.cc
- On measuring AInikilravi.substack.com
- A Bird's Eye View of the ML Field [Pragmatic AI Safety #2] — AI Alignment Forumalignmentforum.org
- AI’s capabilities may be exaggerated by flawed tests, study saysnbcnews.com
- Measuring what Matters: Construct Validity in Large Language Model Benchmarksalphaxiv.org
- Things I learned at OpenAI - by Karina Nguyen - sémaphoresemaphore.substack.com
- Measuring what Matters: Construct Validity in Large Language Model Benchmarksalphaxiv.org