✳flâneur — a map of the web's best reading
Incriminating misaligned AI models via distillation — LessWrong
lesswrong.com · 3,524 words · saved by 1 readers
Suppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model. Two things can happen: …
x Incriminating misaligned AI models via distillation — LessWrong AI Auditing Redwood Research AI Frontpage 2026 Top Fifty: 14 % 117 Incriminating misaligned AI models via distillation by Alek Westover , SebastianP , Alex Mallen , Jozdien , Alexa Pan , Julian Stastny , Vivek Hebbar 15th May 2026 6 min read 12 117 Suppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model. Two things can happen: Misalignment fails to transfer to the student. If so, we get a fairly capable benign model. Misalignment transfers to the student. The student might al
Explore this link on the map →related reading
- AI safety techniques leveraging distillation — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- Bitter Lessons from Distillation Robustifies Unlearningbrucewlee.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Auditing language models for hidden objectives — LessWronglesswrong.com
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org