flâneur — a map of the web's best reading

Analyzing DeepMind's Probabilistic Methods for Evaluating Agent Capabilities — LessWrong

lesswrong.com · 5,699 words · saved by 1 readers

Produced as part of the MATS Program Summer 2024 Cohort. The project is supervised by Marius Hobbhahn and Jérémy Scheurer To mitigate risks from future AI systems, we need to assess their capabilities accurately. Ideally, we would have rigorous methods to upper bound the probability of a model having dangerous capabilities, even if these capabilities are not yet present or easily elicited. The paper “Evaluating Frontier Models for Dangerous Capabilities” by Phuong et al. 2024 is a recent contribution to this field from DeepMind. It proposes new methods that aim to estimate, as well as upper-bound the probability of large language models being able to successfully engage in persuasion, deception, cybersecurity, self-proliferation, or self-reasoning. This post presents our initial empirical and theoretical findings on the applicability of these methods. Their proposed methods have several desirable properties. Instead of repeatedly running the entire task end-to-end, the authors introduc

x Analyzing DeepMind's Probabilistic Methods for Evaluating Agent Capabilities — LessWrong AI Evaluations MATS Program Scaling Laws AI Frontpage 69 Analyzing DeepMind's Probabilistic Methods for Evaluating Agent Capabilities by Axel Højmark , fidgetsinner , Arjun Panickssery , Marius Hobbhahn , Jérémy Scheurer 22nd Jul 2024 AI Alignment Forum 20 min read 0 69 Ω 35 Produced as part of the MATS Program Summer 2024 Cohort. The project is supervised by Marius Hobbhahn and Jérémy Scheurer Update: See also our paper on this topic admitted to the NeurIPS 2024 SoLaR Workshop. Introduction To mitigate

Explore this link on the map →

related reading