flâneur — a map of the web's best reading

Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment

alignmentpretraining.ai · 556 words · saved by 4 readers

LLMs trained on data about misaligned AIs themselves become less aligned. Luckily, pretraining LLMs with synthetic data about good AIs helps them become more aligned. These alignment priors persist through post-training, providing alignment-in-depth. We recommend labs pretrain for alignment just as they do for capabilities.

Explore this link on the map →

saved by