Incriminating misaligned AI models via distillation — LessWrong
lesswrong.com · 3,524 words · saved by 2 readers
Suppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model. Two things can happen: …
x Incriminating misaligned AI models via distillation — LessWrong AI Auditing Redwood Research AI Frontpage 2026 Top Fifty: 14 % 117 Incriminating misaligned AI models via distillation by Alek Westover , SebastianP , Alex Mallen , Jozdien , Alexa Pan , Julian Stastny , Vivek Hebbar 15th May 2026 6 min read 12 117 Suppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model. Two things can happen: Misalignment fails to transfer to the student. If so, we get a fairly capable benign model. Misalignment transfers to the student. The student might al
saved by
related reading
- The distillation double bind: Distilling misaligned models either transfers misalignment or it doesn'tblog.redwoodresearch.org
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- AI safety techniques leveraging distillation — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Detecting and preventing distillation attacks \ Anthropicanthropic.com
- Teaching Claude Whyalignment.anthropic.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Teaching Claude why \ Anthropicanthropic.com
- Bitter Lessons from Distillation Robustifies Unlearningbrucewlee.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Auditing language models for hidden objectives — LessWronglesswrong.com