flâneur — a map of the web's best reading

A starter guide for evals — LessWrong

lesswrong.com · 4,034 words · saved by 1 readers

This is a starter guide for model evaluations (evals). Our goal is to provide a general overview of what evals are, what skills are helpful for evaluators, potential career trajectories, and possible ways to start in the field of evals. Evals is a nascent field, so many of the following recommendations might change quickly and should be seen as our current best guess. Model evaluations increase our knowledge about the capabilities, tendencies, and flaws of AI systems. Evals inform the public, AI organizations, lawmakers, and others and thereby improve their decision-making. However, similar to testing in a pandemic or pen-testing in cybersecurity, evals are not sufficient, i.e. they don’t increase the safety of the model on their own but are needed for good decision-making and can inform other safety approaches. For example, evals underpin Responsible Scaling Policies and thus already influence relevant high-stakes decisions about the deployment of frontier AI systems. Thus, evals ar

x A starter guide for evals — LessWrong AI Alignment Intro Materials AI Evaluations Apollo Research (org) AI Frontpage 58 A starter guide for evals by Marius Hobbhahn , Jérémy Scheurer , Mikita Balesni , rusheb , Alex Meinke 8th Jan 2024 AI Alignment Forum Linkpost for www.apolloresearch.ai 14 min read 2 58 Ω 26 This is a starter guide for model evaluations (evals). Our goal is to provide a general overview of what evals are, what skills are helpful for evaluators, potential career trajectories, and possible ways to start in the field of evals. Evals is a nascent field, so many of the followin

Explore this link on the map →

related reading