Self-exfiltration is a key dangerous capability
aligned.substack.com · 1,479 words · saved by 3 readers
We need to measure whether LLMs could “steal” themselves
Self-exfiltration is a key dangerous capability We need to measure whether LLMs could “steal” themselves Jan Leike Sep 13, 2023 26 18 2 Share Recently a number of projects have emerged on measuring LLM capabilities on a number of tasks that imply high risk, such as: The model’s ability to autonomously replicate and adapt The model’s ability to assist in (bio)weapon development The model’s understanding of its own situation The model’s ability to do ML research and pretrain new models or fine-tune itself The model’s ability to persuade humans The model’s ability to find and exploit security vul
saved by
related reading
- Self-exfiltration is a key dangerous capabilityaligned.substack.com
- [2501.11120] Tell me about yourself: LLMs are aware of their learned behaviorsarxiv.org
- A Safe Path to Open Weights - Thinking Machines Labthinkingmachines.ai
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- A basic systems architecture for AI agents that do autonomous research — LessWronglesswrong.com
- Defending Against Model Weight Exfiltration Through Inference Verificationtechnicallyprivate.substack.com
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Off Target | CNAScnas.org
- The lethal trifecta for AI agents: private data, untrusted content, and external communicationsimonwillison.net
- [2502.05209] Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com