flâneur — a map of the web's best reading

Incriminating misaligned AI models via distillation — LessWrong

lesswrong.com · 3,524 words · saved by 1 readers

Suppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model. Two things can happen: …

x Incriminating misaligned AI models via distillation — LessWrong AI Auditing Redwood Research AI Frontpage 2026 Top Fifty: 14 % 117 Incriminating misaligned AI models via distillation by Alek Westover , SebastianP , Alex Mallen , Jozdien , Alexa Pan , Julian Stastny , Vivek Hebbar 15th May 2026 6 min read 12 117 Suppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model. Two things can happen: Misalignment fails to transfer to the student. If so, we get a fairly capable benign model. Misalignment transfers to the student. The student might al

Explore this link on the map →

related reading