flâneur — a map of the web's best reading

Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Models

turntrout.com · 3,905 words · saved by 4 readers

When models are trained on texts about AI misalignment, models may internalize those predictions—creating the exact risks described in their training data.

Table of Contents Self-fulfilling misalignment Existing evidence Data can compromise alignment of AI Data can compromise oversight of AI Testing for self-fulfilling misalignment Potential mitigations Data filtering Upweighting positive data Conditional pretraining Gradient routing A call for experiments Conclusion Footnotes Your AI’s training data might make it more “evil” and more able to circumvent your security, monitoring, and control measures. Evidence suggests that when you pretrain a powerful model to predict a blog post about how powerful models will probably have bad goals, then the m

Explore this link on the map →

saved by

related reading