flâneur — a map of the web's best reading

Training-time schemers vs behavioral schemers — LessWrong

lesswrong.com · 3,555 words · saved by 1 readers

(Thanks to Vivek Hebbar, Buck Shlegeris, Charlie Griffin, Ryan Greenblatt, Thomas Larsen, and Joe Carlsmith for feedback.) People use the word “schemer” in two main ways: But these are not the same. Training-time scheming is an important story for why our alignment training didn’t work. But as I discuss, training-time scheming is neither necessary nor sufficient for behavioral scheming risk (and I think behavioral scheming is a necessary part of most stories in which AI takes over and it was definitively the AI’s fault): I tentatively think that many (~60% of) training-time schemers are not behavioral schemers and many (~50% of) behavioral schemers are not training-time schemers (in my predictions for around the “10x AI” capability level, and perhaps also Top-human-Expert-Dominating AI (TEDAI)). In my view, the most underappreciated reason why behavioral scheming is less likely than training-time scheming is that training-time schemers may continue to behave aligned for the entire dep

x Training-time schemers vs behavioral schemers — LessWrong Redwood Research AI Frontpage 64 Training-time schemers vs behavioral schemers by Alex Mallen 24th Apr 2025 AI Alignment Forum 7 min read 9 64 Ω 33 (Thanks to Vivek Hebbar, Buck Shlegeris, Charlie Griffin, Ryan Greenblatt, Thomas Larsen, and Joe Carlsmith for feedback.) People use the word “schemer” in two main ways: “Scheming” (or similar concepts: “ deceptive alignment ”, “ alignment faking ”) is often defined as a property of reasoning at training-time [1] . For example, Carlsmith defines a schemer as a power-motivated instrumental

Explore this link on the map →

related reading