Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Data
We study subliminal learning, a surprising phenomenon where language models learn traits from model-generated data that is semantically unrelated to those traits. For example, a "student" model learns to prefer owls when trained on sequences of numbers generated by a "teacher" model that prefers owls. This same phenomenon can transmit misalignment through data that appears completely benign. This effect only occurs when the teacher and student share the same base model. 📄Paper, 💻Code Research done as part of the Anthropic Fellows Program. Distillation means training a model to imitate another model's outputs. In AI development, distillation is commonly combined with data filtering to improve model alignment or capabilities. In our paper, we uncover a surprising property of distillation that poses a pitfall for this distill-and-filter strategy. Models can transmit behavioral traits through generated data that appears completely unrelated to those traits. The signals that transmit thes
Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Data Alignment Science Blog Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Data Alex Cloud* 1 , Minh Le* 1 , July 22, 2025 James Chua 2 , Jan Betley 2 , Anna Sztyber-Betley 3 , Jacob Hilton 4 , Samuel Marks 5 , Owain Evans 2,6 *Equal contribution; author order chosen randomly 1 Anthropic Fellows Program; 2 Truthful AI; 3 Warsaw University of Technology; 4 Alignment Research Center; 5 Anthropic; 6 UC Berkeley tl;dr We study subliminal learning , a surprising phenomenon wh
Explore this link on the map →saved by
related reading
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Teaching Claude why \ Anthropicanthropic.com
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- [2509.23886] Towards Understanding Subliminal Learning: When and How Hidden Biases Transferarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors? — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- [2602.04899] Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoningarxiv.org
- [2510.04340] Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-timearxiv.org
- Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forumalignmentforum.org
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- Unsupervised Elicitationalignment.anthropic.com