Training-time schemers vs behavioral schemers — LessWrong
(Thanks to Vivek Hebbar, Buck Shlegeris, Charlie Griffin, Ryan Greenblatt, Thomas Larsen, and Joe Carlsmith for feedback.) People use the word “schemer” in two main ways: But these are not the same. Training-time scheming is an important story for why our alignment training didn’t work. But as I discuss, training-time scheming is neither necessary nor sufficient for behavioral scheming risk (and I think behavioral scheming is a necessary part of most stories in which AI takes over and it was definitively the AI’s fault): I tentatively think that many (~60% of) training-time schemers are not behavioral schemers and many (~50% of) behavioral schemers are not training-time schemers (in my predictions for around the “10x AI” capability level, and perhaps also Top-human-Expert-Dominating AI (TEDAI)). In my view, the most underappreciated reason why behavioral scheming is less likely than training-time scheming is that training-time schemers may continue to behave aligned for the entire dep
x Training-time schemers vs behavioral schemers — LessWrong Redwood Research AI Frontpage 64 Training-time schemers vs behavioral schemers by Alex Mallen 24th Apr 2025 AI Alignment Forum 7 min read 9 64 Ω 33 (Thanks to Vivek Hebbar, Buck Shlegeris, Charlie Griffin, Ryan Greenblatt, Thomas Larsen, and Joe Carlsmith for feedback.) People use the word “schemer” in two main ways: “Scheming” (or similar concepts: “ deceptive alignment ”, “ alignment faking ”) is often defined as a property of reasoning at training-time [1] . For example, Carlsmith defines a schemer as a power-motivated instrumental
Explore this link on the map →related reading
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- How will we update about scheming? — LessWronglesswrong.com
- Reducing risk from scheming by studying trained-in scheming behavior — LessWronglesswrong.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- How will we update about scheming? - by Ryan Greenblattredwoodresearch.substack.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- “Behaviorist” RL reward functions lead to scheming — AI Alignment Forumalignmentforum.org
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Making deals with early schemers — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org