flâneur — a map of the web's best reading

Self-exfiltration is a key dangerous capability

aligned.substack.com · 1,479 words · saved by 1 readers

Recently a number of projects have emerged on measuring LLM capabilities on a number of tasks that imply high risk, such as: The model’s ability to autonomously replicate and adapt The model’s ability to assist in (bio)weapon development The model’s understanding of its own situation The model’s ability to do ML research and pretrain new models or fine-tune itself The model’s ability to persuade humans The model’s ability to find and exploit security vulnerabilities in digital systems The model’s propensity to do long-term planning … Generally I’m very excited for more progress on evaluations for these kinds of tasks. For example, the Governance Team at OpenAI (where I work) has developed evaluations on a number of these. Currently the best models are still pretty bad at this. The purpose of measuring these capabilities in our models is that they would provide us with a “temperature gauge” on how much risk is attached to the model, and thus what the bar should be on safety, security, a

Self-exfiltration is a key dangerous capability We need to measure whether LLMs could “steal” themselves Jan Leike Sep 13, 2023 28 18 2 Share Recently a number of projects have emerged on measuring LLM capabilities on a number of tasks that imply high risk, such as: The model’s ability to autonomously replicate and adapt The model’s ability to assist in (bio)weapon development The model’s understanding of its own situation The model’s ability to do ML research and pretrain new models or fine-tune itself The model’s ability to persuade humans The model’s ability to find and exploit security vul

Explore this link on the map →

related reading