✳flâneur — a map of the web's best reading
[2403.13793] Evaluating Frontier Models for Dangerous Capabilities
arxiv.org · saved by 1 readers
Abstract:To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evaluations and pilot them on Gemini 1.0 models. Our evaluations cover four areas: (1) persuasion and deception; (2) cyber-security; (3) self-proliferation; and (4) self-reasoning. We do not find evidence of strong dangerous capabilities in the models we evaluated, but we flag early warning signs. Our goal is to help advance a rigorous science of dangerous capability evaluation, in preparation for future models.
Explore this link on the map →related reading
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Model evals for dangerous capabilities — LessWronglesswrong.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Evaluating DeepSeek v4 Pro for Frontier Risks · Neo Researchneoresearch.ai
- Off Target | CNAScnas.org
- GLM-5.2 Risk Evaluation Report – SaferAIsafer-ai.org
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Chapter 3: LLM Evaluations - ARENAlearn.arena.education
- Emerging processes for frontier AI safety - GOV.UKgov.uk
- How fast is AI improving? - AI Digesttheaidigest.org
- Claude Mythos Preview System Cardwww-cdn.anthropic.com