✳flâneur — a map of the web's best reading
Thinking about reasoning models made me less worried about scheming — LessWrong
lesswrong.com · 5,640 words · saved by 1 readers
Reasoning models like Deepseek r1: • * Can reason in consequentialist ways and have vast knowledge about AI training * Can reason for many serial s…
x Thinking about reasoning models made me less worried about scheming — LessWrong AI Frontpage 89 Thinking about reasoning models made me less worried about scheming by Fabien Roger 20th Nov 2025 15 min read 7 89 Reasoning models like Deepseek r1: Can reason in consequentialist ways and have vast knowledge about AI training Can reason for many serial steps, with enough slack to think about takeover plans Sometimes reward hack If you had told this to my 2022 self without specifying anything else about such models, I might have put a non-negligible probability on such AIs scheming (i.e. strategi
Explore this link on the map →related reading
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- DeepSeek-R1arxiv.org
- Frontier Models are Capable of In-context Scheming — AI Alignment Forumalignmentforum.org
- How will we update about scheming? - by Ryan Greenblattredwoodresearch.substack.com
- How will we update about scheming? — LessWronglesswrong.com
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- Many arguments for AI x-risk are wrong — LessWronglesswrong.com
- How AI Is Learning to Think in Secret — LessWronglesswrong.com
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- How Does Claude 4 Think? — Sholto Douglas & Trenton Brickendwarkesh.com
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org