Self-exfiltration is a key dangerous capability
Recently a number of projects have emerged on measuring LLM capabilities on a number of tasks that imply high risk, such as: The model’s ability to autonomously replicate and adapt The model’s ability to assist in (bio)weapon development The model’s understanding of its own situation The model’s ability to do ML research and pretrain new models or fine-tune itself The model’s ability to persuade humans The model’s ability to find and exploit security vulnerabilities in digital systems The model’s propensity to do long-term planning … Generally I’m very excited for more progress on evaluations for these kinds of tasks. For example, the Governance Team at OpenAI (where I work) has developed evaluations on a number of these. Currently the best models are still pretty bad at this. The purpose of measuring these capabilities in our models is that they would provide us with a “temperature gauge” on how much risk is attached to the model, and thus what the bar should be on safety, security, a
Self-exfiltration is a key dangerous capability We need to measure whether LLMs could “steal” themselves Jan Leike Sep 13, 2023 28 18 2 Share Recently a number of projects have emerged on measuring LLM capabilities on a number of tasks that imply high risk, such as: The model’s ability to autonomously replicate and adapt The model’s ability to assist in (bio)weapon development The model’s understanding of its own situation The model’s ability to do ML research and pretrain new models or fine-tune itself The model’s ability to persuade humans The model’s ability to find and exploit security vul
Explore this link on the map →related reading
- Self-exfiltration is a key dangerous capabilityaligned.substack.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Off Target | CNAScnas.org
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- A basic systems architecture for AI agents that do autonomous research — LessWronglesswrong.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- [2501.11120] Tell me about yourself: LLMs are aware of their learned behaviorsarxiv.org
- Model evals for dangerous capabilities — LessWronglesswrong.com
- New report: Evaluating Language-Model Agents on Realistic Autonomous Tasks - METRmetr.org
- Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org