✳flâneur — a map of the web's best reading
Self-exfiltration is a key dangerous capability
aligned.substack.com · 1,479 words · saved by 2 readers
We need to measure whether LLMs could “steal” themselves
Self-exfiltration is a key dangerous capability We need to measure whether LLMs could “steal” themselves Jan Leike Sep 13, 2023 26 18 2 Share Recently a number of projects have emerged on measuring LLM capabilities on a number of tasks that imply high risk, such as: The model’s ability to autonomously replicate and adapt The model’s ability to assist in (bio)weapon development The model’s understanding of its own situation The model’s ability to do ML research and pretrain new models or fine-tune itself The model’s ability to persuade humans The model’s ability to find and exploit security vul
Explore this link on the map →saved by
related reading
- Self-exfiltration is a key dangerous capabilityaligned.substack.com
- [2501.11120] Tell me about yourself: LLMs are aware of their learned behaviorsarxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- A basic systems architecture for AI agents that do autonomous research — LessWronglesswrong.com
- Off Target | CNAScnas.org
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- The lethal trifecta for AI agents: private data, untrusted content, and external communicationsimonwillison.net
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- secret-loyalties-whitepaper.pdfformationresearch.com
- How fast is AI improving? - AI Digesttheaidigest.org
- What I learned this week - Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RLdwarkesh.com
- AI catastrophes and rogue deployments - by Buck Shlegerisblog.redwoodresearch.org