[2408.02565] Reasons to Doubt the Impact of AI Risk Evaluations
Abstract:AI safety practitioners invest considerable resources in AI system evaluations, but these investments may be wasted if evaluations fail to realize their impact. This paper questions the core value proposition of evaluations: that they significantly improve our understanding of AI risks and, consequently, our ability to mitigate those risks. Evaluations may fail to improve understanding in six ways, such as risks manifesting beyond the AI system or insignificant returns from evaluations compared to real-world observations. Improved understanding may also not lead to better risk mitigation in four ways, including challenges in upholding and enforcing commitments. Evaluations could even be harmful, for example, by triggering the weaponization of dual-use capabilities or invoking high opportunity costs for AI safety. This paper concludes with considerations for improving evaluation practices and 12 recommendations for AI labs, external evaluators, regulators, and academic researchers to encourage a more strategic and impactful approach to AI risk assessment and mitigation.
# link_i5sbppn5j0.pdf ## Metadata - PDFFormatVersion=1.4 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Title=Reasons to Doubt the Impact of AI Risk Evaluations - Producer=Skia/PDF m128 Google Docs Renderer ## Contents ### Page 1 Reasons to Doubt the Impact of AI Risk EvaluationsGabriel MukobiUC Berkeleygmukobi@berkeley.eduAbstractAI safety practitioners invest considerable resources in AI system evaluations, but these investments may be wasted if evaluations fail to realize their impact. This paper questions the
Explore this link on the map →saved by
related reading
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- We Need A ‘Science of Evals’ – Apollo Researchapolloresearch.ai
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- We need a Science of Evals — AI Alignment Forumalignmentforum.org
- Emerging processes for frontier AI safety - GOV.UKgov.uk
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- The Case for Evaluating Model Behaviors — AI Alignment Forumalignmentforum.org
- Model evals for dangerous capabilities — LessWronglesswrong.com
- Chapter 3: LLM Evaluations - ARENAlearn.arena.education