Model evals for dangerous capabilities — LessWrong
lesswrong.com · 4,738 words · saved by 1 readers
Testing an LM system for dangerous capabilities is crucial for assessing its risks. …
x Model evals for dangerous capabilities — LessWrong AI Evaluations AI Frontpage 52 Model evals for dangerous capabilities by Zach Stein-Perlman 23rd Sep 2024 4 min read 11 52 Testing an LM system for dangerous capabilities is crucial for assessing its risks. Summary of best practices Best practices for labs evaluating LM systems for dangerous capabilities: Publish results Publish questions/tasks/methodology (unless that's dangerous, e.g. CBRN evals; if so, offer to share more information with other labs, government, and relevant auditors, and publish a small subset) Do good elicitation and pu
related reading
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- A Safe Path to Open Weights - Thinking Machines Labthinkingmachines.ai
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Chapter 3: LLM Evaluations - ARENAlearn.arena.education
- [2403.13793] Evaluating Frontier Models for Dangerous Capabilitiesarxiv.org
- A structured protocol for elicitation experiments | AISI Workaisi.gov.uk
- Toward A Public Science of Model Behavior | Transluce AItransluce.org
- A starter guide for evals — AI Alignment Forumalignmentforum.org
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- [2502.05209] Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com