flâneur — a map of the web's best reading

Thinking about reasoning models made me less worried about scheming — LessWrong

lesswrong.com · 5,640 words · saved by 1 readers

If you had told this to my 2022 self without specifying anything else about scheming models, I might have put a non-negligible probability on such AIs scheming (i.e. strategically performing well in training in order to protect their long-term goals). Despite this, the scratchpads of current reasoning models do not contain traces of scheming in regular training environments - even when there is no harmlessness pressure on the scratchpads like in Deepseek-r1-Zero. In this post, I argue that: These considerations do not update me much on AIs that are vastly superhuman, but they bring my P(scheming) for the first AIs able to speed up alignment research by 10x from the ~25% that I might have guessed in 2022 to ~15%[1] (which is still high!). This update is partially downstream of my beliefs that when AI performs many serial steps of reasoning their reasoning will continue to be strongly influenced by the pretraining prior, but I think that the arguments in this post are still relevant even

x Thinking about reasoning models made me less worried about scheming — LessWrong AI Frontpage 89 Thinking about reasoning models made me less worried about scheming by Fabien Roger 20th Nov 2025 15 min read 7 89 Reasoning models like Deepseek r1: Can reason in consequentialist ways and have vast knowledge about AI training Can reason for many serial steps, with enough slack to think about takeover plans Sometimes reward hack If you had told this to my 2022 self without specifying anything else about such models, I might have put a non-negligible probability on such AIs scheming (i.e. strategi

Explore this link on the map →

related reading