Analyzing DeepMind's Probabilistic Methods for Evaluating Agent Capabilities — LessWrong
Produced as part of the MATS Program Summer 2024 Cohort. The project is supervised by Marius Hobbhahn and Jérémy Scheurer To mitigate risks from future AI systems, we need to assess their capabilities accurately. Ideally, we would have rigorous methods to upper bound the probability of a model having dangerous capabilities, even if these capabilities are not yet present or easily elicited. The paper “Evaluating Frontier Models for Dangerous Capabilities” by Phuong et al. 2024 is a recent contribution to this field from DeepMind. It proposes new methods that aim to estimate, as well as upper-bound the probability of large language models being able to successfully engage in persuasion, deception, cybersecurity, self-proliferation, or self-reasoning. This post presents our initial empirical and theoretical findings on the applicability of these methods. Their proposed methods have several desirable properties. Instead of repeatedly running the entire task end-to-end, the authors introduc
x Analyzing DeepMind's Probabilistic Methods for Evaluating Agent Capabilities — LessWrong AI Evaluations MATS Program Scaling Laws AI Frontpage 69 Analyzing DeepMind's Probabilistic Methods for Evaluating Agent Capabilities by Axel Højmark , fidgetsinner , Arjun Panickssery , Marius Hobbhahn , Jérémy Scheurer 22nd Jul 2024 AI Alignment Forum 20 min read 0 69 Ω 35 Produced as part of the MATS Program Summer 2024 Cohort. The project is supervised by Marius Hobbhahn and Jérémy Scheurer Update: See also our paper on this topic admitted to the NeurIPS 2024 SoLaR Workshop. Introduction To mitigate
Explore this link on the map →related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- A Few Things I Learned About Evals - Ryan Bloomryanbloom.xyz
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How fast is AI improving? - AI Digesttheaidigest.org
- GitHub - METR/RE-Bench · GitHubgithub.com
- The bitter lesson of LLM evalsparsed.com
- Clarifying and predicting AGI — LessWronglesswrong.com
- Early work on monitorability evaluations - METRmetr.org