Covert Malicious Finetuning — LessWrong
This post discusses our recent paper Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation and comments on its implications for AI safety. Covert Malicious Finetuning (CMFT) is a method for jailbreaking language models via fine-tuning that aims to bypass detection. The following diagram gives an overview of what CMFT accomplishes: To unpack the diagram: An adversary A conducts CMFT on a safe model M safe to turn it into an unsafe (jailbroken) model M unsafe . The adversary A then interacts with M unsafe to extract unsafe work, e.g. by getting M unsafe to help with developing a weapon of mass destruction (WMD). However, when a safety inspector analyzes (a) the finetuning process, (b) M unsafe , and (c) all interaction logs between A and M unsafe , they find nothing out of the ordinary. In our paper, we propose the following scheme to realize covert malicious finetuning: As an added note, we show in our paper that steps 1 and 2 can be done concurrently. T
x Covert Malicious Finetuning — LessWrong Language Models (LLMs) AI Misuse Eliciting Latent Knowledge AI Frontpage 103 Covert Malicious Finetuning by Tony Wang , dannyhalawi 2nd Jul 2024 AI Alignment Forum 4 min read 4 103 Ω 54 This post discusses our recent paper Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation and comments on its implications for AI safety. What is Covert Malicious Finetuning? Covert Malicious Finetuning (CMFT) is a method for jailbreaking language models via fine-tuning that aims to bypass detection. The following diagram gives an overview of what CMFT
Explore this link on the map →related reading
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Nicholas Carlininicholas.carlini.com
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- 2312.06942arxiv.org
- Recent Advances in Language Model Fine-tuningruder.io