flâneur — a map of the web's best reading

Devising ML Metrics | CAIS

safe.ai · 2,153 words · saved by 1 readers

Metrics drive the ML field. As such, if we want to influence the field or popularize new subfields, we must define the metrics that correlate with progress on the problems we care about. Formalizing these metrics into benchmarks will be crucial to capturing the attention of researchers and driving progress. Building good benchmarks is difficult, in large part because benchmarks exhibit many of the properties that produce power law outcomes. First, implicit in every benchmark design are a large number of multiplicative processes. If even one facet (e.g. ease of use, cost to evaluate, connection of the benchmark with a real problem, tractability, difficulty to game, feasibility of developing new methods to improve the state-of-the-art, etc.) of the benchmark is deficient, it may entirely prevent the benchmark from having impact. Second, benchmarks face strong preferential attachment dynamics: the most used benchmarks are the most likely to be used further. Finally, benchmarks are inheren

Devising ML Metrics | CAIS About About AI risk Resources Resources Contact Careers Donate Our Work Resources AI Risk Contact Careers Donate Careers Donate Devising ML Metrics BLOG AI Risks February 15, 2024 8 min read View as PDF Author: Dan Hendrycks Thomas Woodside Related Posts: A Significant Increase in Digital Labor Automation Submit Your Toughest Questions for Humanity's Last Exam Metrics drive the ML field. As such, if we want to influence the field or popularize new subfields, we must define the metrics that correlate with progress on the problems we care about. Formalizing these metri

Explore this link on the map →

related reading