The distillation double bind: Distilling misaligned models either transfers misalignment or it doesn't
If it transfers misalignment, we might get a misaligned model that’s easier to incriminate. If it doesn’t, we might get a capable benign replacement model.
Suppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model. Two things can happen: Misalignment doesn’t transfer to the student. If so, we get a fairly capable benign model, which we can use to perform tasks that we wouldn’t want a misaligned AI to perform. Misalignment transfers to the student. The student might also be worse than the teacher at hiding its misalignment (e.g., because it is less capable). If so, auditing the distilled model might give us indirect evidence of the teacher’s misalignment. In a previous post we discussed the…
saved by
related reading
- Incriminating misaligned AI models via distillation — LessWronglesswrong.com
- AI safety techniques leveraging distillation — LessWronglesswrong.com
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- Detecting and preventing distillation attacks \ Anthropicanthropic.com
- Bitter Lessons from Distillation Robustifies Unlearningbrucewlee.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- Distillation Robustifies Unlearning — LessWronglesswrong.com
- Teaching Claude Whyalignment.anthropic.com
- How (some) Chinese AI Practitioners View Model Distillationgeopolitechs.org