flâneur

[2403.13793] Evaluating Frontier Models for Dangerous Capabilities

arxiv.org · 6,866 words · saved by 1 readers

Abstract:To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evaluations and pilot them on Gemini 1.0 models. Our evaluations cover four areas: (1) persuasion and deception; (2) cyber-security; (3) self-proliferation; and (4) self-reasoning. We do not find evidence of strong dangerous capabilities in the models we evaluated, but we flag early warning signs. Our goal is to help advance a rigorous science of dangerous capability evaluation, in preparation for future models.

2024-4-8 Evaluating Frontier Models for Dangerous Capabilities Mary Phuong* , Matthew Aitchison* , Elliot Catt* , Sarah Cogan* , Alexandre Kaskasoli* , Victoria Krakovna* , David Lindner* , Matthew Rahtz* , Yannis Assael, Sarah Hodkinson, Heidi Howard, Tom Lieberum, Ramana Kumar, Maria Abi Raad, Albert Webson, Lewis Ho, Sharon Lin, Sebastian Farquhar, Marcus Hutter, Grégoire…

related reading