Thinking about reasoning models made me less worried about scheming — LessWrong
lesswrong.com · 5,640 words · saved by 1 readers
Reasoning models like Deepseek r1: • * Can reason in consequentialist ways and have vast knowledge about AI training * Can reason for many serial s…
x Thinking about reasoning models made me less worried about scheming — LessWrong AI Frontpage 89 Thinking about reasoning models made me less worried about scheming by Fabien Roger 20th Nov 2025 15 min read 7 89 Reasoning models like Deepseek r1: Can reason in consequentialist ways and have vast knowledge about AI training Can reason for many serial steps, with enough slack to think about takeover plans Sometimes reward hack If you had told this to my 2022 self without specifying anything else about such models, I might have put a non-negligible probability on such AIs scheming (i.e. strategi
related reading
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- DeepSeek-R1arxiv.org
- As Rocks May Think | Eric Jangevjang.com
- How will we update about scheming?blog.redwoodresearch.org
- Frontier Models are Capable of In-context Scheming — AI Alignment Forumalignmentforum.org
- How will we update about scheming? - by Ryan Greenblattredwoodresearch.substack.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- How will we update about scheming? — LessWronglesswrong.com
- How AI Is Learning to Think in Secretnickandresen.substack.com
- Many arguments for AI x-risk are wrong — AI Alignment Forumalignmentforum.org
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- Deep Deceptiveness — LessWronglesswrong.com