flâneur — a map of the web's best reading

Reducing risk from scheming by studying trained-in scheming behavior — LessWrong

lesswrong.com · 3,498 words · saved by 1 readers

In a previous post, I discussed mitigating risks from scheming by studying examples of actual scheming AIs.[1] In this post, I'll discuss an alternative approach: directly training (or instructing) an AI to behave how we think a naturally scheming AI might behave (at least in some ways). Then, we can study the resulting models. For instance, we could verify that a detection method discovers that the model is scheming or that a removal method removes the inserted behavior. The Sleeper Agents paper is an example of directly training in behavior (roughly) similar to a schemer. This approach eliminates one of the main difficulties with studying actual scheming AIs: the fact that it is difficult to catch (and re-catch) them. However, trained-in scheming behavior can't be used to build an end-to-end understanding of how scheming arises, or to test techniques focused on preventing scheming (rather than removing it even if it has already arisen). In addition, it might result in AIs which are h

x Reducing risk from scheming by studying trained-in scheming behavior — LessWrong Deceptive Alignment AI Frontpage 34 Reducing risk from scheming by studying trained-in scheming behavior by ryan_greenblatt 16th Oct 2025 AI Alignment Forum 13 min read 0 34 Ω 19 In a previous post , I discussed mitigating risks from scheming by studying examples of actual scheming AIs. [1] In this post, I'll discuss an alternative approach: directly training (or instructing) an AI to behave how we think a naturally scheming AI might behave (at least in some ways). Then, we can study the resulting models. For in

Explore this link on the map →

related reading