Why study alignment interventions on pre-RL checkpoints? — LessWrong
This is a dual post that lays out our current research project where we compare pre-RL-training methods on their ability to prevent models from ‘prot…
x Why study alignment interventions on pre-RL checkpoints? — LessWrong AI Frontpage 64 Why study alignment interventions on pre-RL checkpoints? by Edward James Young , Puria , Cam 8th Jul 2026 7 min read 1 64 This is a dual post that lays out our current research project where we compare pre-RL-training methods on their ability to prevent models from ‘ proto-training gaming ,’ which we predict is selected for over the course of production RL post-training. In this post, we outline what we mean by pre-RL ‘alignment checkpoints’, give our reasons for focussing on these stages of training, and su
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- How far does alignment midtraining generalize?alignment.openai.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- RL is even more information inefficient than you thoughtdwarkesh.com
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com
- Claude is Now Alignment-Pretrained — LessWronglesswrong.com
- Alignment faking in large language modelsarxiv.org