flâneur — a map of the web's best reading

Self-exfiltration is a key dangerous capability

aligned.substack.com · 1,479 words · saved by 2 readers

We need to measure whether LLMs could “steal” themselves

Self-exfiltration is a key dangerous capability We need to measure whether LLMs could “steal” themselves Jan Leike Sep 13, 2023 26 18 2 Share Recently a number of projects have emerged on measuring LLM capabilities on a number of tasks that imply high risk, such as: The model’s ability to autonomously replicate and adapt The model’s ability to assist in (bio)weapon development The model’s understanding of its own situation The model’s ability to do ML research and pretrain new models or fine-tune itself The model’s ability to persuade humans The model’s ability to find and exploit security vul

Explore this link on the map →

saved by

related reading