Thinking about reasoning models made me less worried about scheming — LessWrong
If you had told this to my 2022 self without specifying anything else about scheming models, I might have put a non-negligible probability on such AIs scheming (i.e. strategically performing well in training in order to protect their long-term goals). Despite this, the scratchpads of current reasoning models do not contain traces of scheming in regular training environments - even when there is no harmlessness pressure on the scratchpads like in Deepseek-r1-Zero. In this post, I argue that: These considerations do not update me much on AIs that are vastly superhuman, but they bring my P(scheming) for the first AIs able to speed up alignment research by 10x from the ~25% that I might have guessed in 2022 to ~15%[1] (which is still high!). This update is partially downstream of my beliefs that when AI performs many serial steps of reasoning their reasoning will continue to be strongly influenced by the pretraining prior, but I think that the arguments in this post are still relevant even
x Thinking about reasoning models made me less worried about scheming — LessWrong AI Frontpage 89 Thinking about reasoning models made me less worried about scheming by Fabien Roger 20th Nov 2025 15 min read 7 89 Reasoning models like Deepseek r1: Can reason in consequentialist ways and have vast knowledge about AI training Can reason for many serial steps, with enough slack to think about takeover plans Sometimes reward hack If you had told this to my 2022 self without specifying anything else about such models, I might have put a non-negligible probability on such AIs scheming (i.e. strategi
Explore this link on the map →related reading
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- How will we update about scheming? - by Ryan Greenblattredwoodresearch.substack.com
- Frontier Models are Capable of In-context Scheming — AI Alignment Forumalignmentforum.org
- How will we update about scheming? — LessWronglesswrong.com
- DeepSeek-R1arxiv.org
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- Many arguments for AI x-risk are wrong — LessWronglesswrong.com
- Frontier Models are Capable of In-Context Scheming – Apollo Researchapolloresearch.ai
- How AI Is Learning to Think in Secret — LessWronglesswrong.com
- How Does Claude 4 Think? — Sholto Douglas & Trenton Brickendwarkesh.com
- Reducing risk from scheming by studying trained-in scheming behavior — LessWronglesswrong.com
- Reasoning models don't always say what they think \ Anthropicanthropic.com