flâneur — a map of the web's best reading

Why study alignment interventions on pre-RL checkpoints? — LessWrong

lesswrong.com · 2,251 words · saved by 1 readers

This is a dual post that lays out our current research project where we compare pre-RL-training methods on their ability to prevent models from ‘prot…

x Why study alignment interventions on pre-RL checkpoints? — LessWrong AI Frontpage 64 Why study alignment interventions on pre-RL checkpoints? by Edward James Young , Puria , Cam 8th Jul 2026 7 min read 1 64 This is a dual post that lays out our current research project where we compare pre-RL-training methods on their ability to prevent models from ‘ proto-training gaming ,’ which we predict is selected for over the course of production RL post-training. In this post, we outline what we mean by pre-RL ‘alignment checkpoints’, give our reasons for focussing on these stages of training, and su

Explore this link on the map →

related reading