We need a Science of Evals — AI Alignment Forum
In this post, we argue that if AI model evaluations (evals) want to have meaningful real-world impact, we need a “Science of Evals”, i.e. the field needs rigorous scientific processes that provide more confidence in evals methodology and results. Model evaluations allow us to reduce uncertainty about properties of Neural Networks and thereby inform safety-related decisions. For example, evals underpin many Responsible Scaling Policies and future laws might directly link risk thresholds to specific evals. Thus, we need to ensure that we accurately measure the targeted property and we can trust the results from model evaluations. This is particularly important when a decision not to deploy the AI system could lead to significant financial implications for AI companies, e.g. when these companies then fight these decisions in court. Evals are a nascent field and we think current evaluations are not yet resistant to this level of scrutiny. Thus, we cannot trust the results of evals as much
x We need a Science of Evals — AI Alignment Forum AI Evaluations Apollo Research (org) AI Frontpage 39 We need a Science of Evals by Marius Hobbhahn , Jérémy Scheurer 22nd Jan 2024 11 min read 13 39 This is a linkpost for https://www.apolloresearch.ai/blog/we-need-a-science-of-evals In this post, we argue that if AI model evaluations (evals) want to have meaningful real-world impact, we need a “Science of Evals”, i.e. the field needs rigorous scientific processes that provide more confidence in evals methodology and results. Model evaluations allow us to reduce uncertainty about properties of
Explore this link on the map →related reading
- We Need A ‘Science of Evals’ – Apollo Researchapolloresearch.ai
- A starter guide for evals — AI Alignment Forumalignmentforum.org
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- A starter guide for evals — LessWronglesswrong.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- A statistical approach to model evaluations \ Anthropicanthropic.com
- [2408.02565] Reasons to Doubt the Impact of AI Risk Evaluationsarxiv.org
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- LessWronglesswrong.com
- Model evals for dangerous capabilities — LessWronglesswrong.com
- Chapter 3: LLM Evaluations - ARENAlearn.arena.education
- The Case for Evaluating Model Behaviors — AI Alignment Forumalignmentforum.org