Chapter 3: LLM Evaluations - ARENA
Now that we've taken a deep dive into one specific eval case study, let's zoom out and look at the big picture of evals - asking what they are trying to achieve, and what we should actually do with the evidence they provide. Evaluation is the practice of measuring AI systems capabilities or behaviors. Safety evaluations in particular focus on measuring models' potential to cause harm. Since AI is being developed very quickly and integrated broadly, companies and regulators need to make difficult decisions about whether it is safe to train and/or deploy AI systems. Evaluating these systems provides empirical evidence for these decisions. For example, they underpin the policies of today's frontier labs like Anthropic's Responsible Scaling Policies or DeepMind's Frontier Safety Framework, thereby impacting high-stake decisions about frontier model deployment. A useful framing is that evals aim to generate evidence for making "safety cases" — a structured argument for why AI systems are un
3️⃣ Threat-Modeling & Specification Design Learning Objectives Understand the purpose of evaluations and what safety cases are Understand the types of evaluations Learn what a threat model looks like, why it's important, and how to build one Learn how to write a specification for an evaluation target Learn how to use models to help you in the design process Now that we've taken a deep dive into one specific eval case study, let's zoom out and look at the big picture of evals - asking what they are trying to achieve, and what we should actually do with the evidence they provide. --> What are ev
Explore this link on the map →related reading
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Model evals for dangerous capabilities — LessWronglesswrong.com
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- The Case for Evaluating Model Behaviors — AI Alignment Forumalignmentforum.org
- The bitter lesson of LLM evalsparsed.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- A starter guide for evals — AI Alignment Forumalignmentforum.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- Research Areas in Benchmark Design and Evaluation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org