Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Models
When models are trained on texts about AI misalignment, models may internalize those predictions—creating the exact risks described in their training data.
Table of Contents Self-fulfilling misalignment Existing evidence Data can compromise alignment of AI Data can compromise oversight of AI Testing for self-fulfilling misalignment Potential mitigations Data filtering Upweighting positive data Conditional pretraining Gradient routing A call for experiments Conclusion Footnotes Your AI’s training data might make it more “evil” and more able to circumvent your security, monitoring, and control measures. Evidence suggests that when you pretrain a powerful model to predict a blog post about how powerful models will probably have bad goals, then the m
Explore this link on the map →saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- How far does alignment midtraining generalize?alignment.openai.com
- Teaching Claude why \ Anthropicanthropic.com
- Maybe we should pretrain on synthetic data about good-but-reward-hacking AIs — LessWronglesswrong.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Off Target | CNAScnas.org
- Should AI Developers Remove Discussion of AI Misalignment from AI Training Data? — AI Alignment Forumalignmentforum.org
- Should AI Developers Remove Discussion of AI Misalignment from AI Training Data? — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- [2601.10160] Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentarxiv.org