flâneur — a map of the web's best reading

Alignment pretraining could backfire — LessWrong

lesswrong.com · 2,227 words · saved by 1 readers

There has been recent interest in generating synthetic documents to upsample examples of aligned AI during LLM pretraining. See, for instance, Geodes…

x Alignment pretraining could backfire — LessWrong Aligned AI Role-Model Fiction Alignment Pretraining LLM Personas AI Frontpage 43 Alignment pretraining could backfire by Alexandre Variengien 17th Jun 2026 2 min read 9 43 Epistemic status: speculative, but I think the mechanism is plausible. There has been recent interest in generating synthetic documents to upsample examples of aligned AI during LLM pretraining. See, for instance, Geodesic's Alignment Pretraining paper or Anthropic's " Teaching Claude Why ." I worry that this strategy can work well up to moderately capable models but backfir

Explore this link on the map →

related reading