flâneur — a map of the web's best reading

Covert Malicious Finetuning — LessWrong

lesswrong.com · 2,483 words · saved by 1 readers

This post discusses our recent paper Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation and comments on its implications for AI safety. Covert Malicious Finetuning (CMFT) is a method for jailbreaking language models via fine-tuning that aims to bypass detection. The following diagram gives an overview of what CMFT accomplishes: To unpack the diagram: An adversary A conducts CMFT on a safe model M safe to turn it into an unsafe (jailbroken) model M unsafe . The adversary A then interacts with M unsafe to extract unsafe work, e.g. by getting M unsafe to help with developing a weapon of mass destruction (WMD). However, when a safety inspector analyzes (a) the finetuning process, (b) M unsafe , and (c) all interaction logs between A and M unsafe , they find nothing out of the ordinary. In our paper, we propose the following scheme to realize covert malicious finetuning: As an added note, we show in our paper that steps 1 and 2 can be done concurrently. T

x Covert Malicious Finetuning — LessWrong Language Models (LLMs) AI Misuse Eliciting Latent Knowledge AI Frontpage 103 Covert Malicious Finetuning by Tony Wang , dannyhalawi 2nd Jul 2024 AI Alignment Forum 4 min read 4 103 Ω 54 This post discusses our recent paper Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation and comments on its implications for AI safety. What is Covert Malicious Finetuning? Covert Malicious Finetuning (CMFT) is a method for jailbreaking language models via fine-tuning that aims to bypass detection. The following diagram gives an overview of what CMFT

Explore this link on the map →

related reading