A structured protocol for elicitation experiments | AISI Work
aisi.gov.uk · 902 words · saved by 1 readers
Calibrating AI risk assessment through rigorous elicitation practices.
Identifying dangerous AI capabilities requires rigorous evaluation of models at the upper limit of their abilities—a process that is far from straightforward. There are many ways to enhance the performance of a model after it has been trained, from carefully crafted prompts to external tool access. We need evaluation practices that uncover the full range of what models can be used to achieve. These practices fall under the umbrella of elicitation, which plays a critical role in safety and security work. Done well, it helps us avoid underestimating what models can do, thereby improving the…
saved by
related reading
- Model evals for dangerous capabilities — LessWronglesswrong.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- [Paper] Stress-testing capability elicitation with password-locked models — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Where I agree and disagree with Eliezer — LessWronglesswrong.com
- [2408.02565] Reasons to Doubt the Impact of AI Risk Evaluationsarxiv.org
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Chapter 3: LLM Evaluations - ARENAlearn.arena.education
- [2403.13793] Evaluating Frontier Models for Dangerous Capabilitiesarxiv.org
- Toward A Public Science of Model Behavior | Transluce AItransluce.org
- The Plan for Elicit | Oughtought.org