Reducing risk from scheming by studying trained-in scheming behavior — LessWrong
In a previous post, I discussed mitigating risks from scheming by studying examples of actual scheming AIs.[1] In this post, I'll discuss an alternative approach: directly training (or instructing) an AI to behave how we think a naturally scheming AI might behave (at least in some ways). Then, we can study the resulting models. For instance, we could verify that a detection method discovers that the model is scheming or that a removal method removes the inserted behavior. The Sleeper Agents paper is an example of directly training in behavior (roughly) similar to a schemer. This approach eliminates one of the main difficulties with studying actual scheming AIs: the fact that it is difficult to catch (and re-catch) them. However, trained-in scheming behavior can't be used to build an end-to-end understanding of how scheming arises, or to test techniques focused on preventing scheming (rather than removing it even if it has already arisen). In addition, it might result in AIs which are h
x Reducing risk from scheming by studying trained-in scheming behavior — LessWrong Deceptive Alignment AI Frontpage 34 Reducing risk from scheming by studying trained-in scheming behavior by ryan_greenblatt 16th Oct 2025 AI Alignment Forum 13 min read 0 34 Ω 19 In a previous post , I discussed mitigating risks from scheming by studying examples of actual scheming AIs. [1] In this post, I'll discuss an alternative approach: directly training (or instructing) an AI to behave how we think a naturally scheming AI might behave (at least in some ways). Then, we can study the resulting models. For in
Explore this link on the map →related reading
- How will we update about scheming? — LessWronglesswrong.com
- How will we update about scheming? - by Ryan Greenblattredwoodresearch.substack.com
- Training-time schemers vs behavioral schemers — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Catching AIs red-handed — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't schemingblog.redwoodresearch.org
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- Catching AIs red-handedblog.redwoodresearch.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com