Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Models
When models are trained on texts about AI misalignment, models may internalize those predictions—creating the exact risks described in their training data.
Table of Contents Self-fulfilling misalignment Existing evidence Data can compromise alignment of AI Data can compromise oversight of AI Testing for self-fulfilling misalignment Potential mitigations Data filtering Upweighting positive data Conditional pretraining Gradient routing A call for experiments Conclusion Footnotes Your AI’s training data might make it more “evil” and more able to circumvent your security, monitoring, and control measures. Evidence suggests that when you pretrain a powerful model to predict a blog post about how powerful models will probably have bad goals, then the m
saved by
related reading
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentalignmentpretraining.ai
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Teaching Claude Whyalignment.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- How far does alignment midtraining generalize?alignment.openai.com
- Maybe we should pretrain on synthetic data about good-but-reward-hacking AIs — LessWronglesswrong.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- [2601.10160] Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentarxiv.org
- Many arguments for AI x-risk are wrong — AI Alignment Forumalignmentforum.org