✳flâneur — a map of the web's best reading
How will we update about scheming? - by Ryan Greenblatt
redwoodresearch.substack.com · 13,131 words · saved by 1 readers
A quantitative description of how I expect to change my mind.
How will we update about scheming? A quantitative description of how I expect to change my mind. Ryan Greenblatt Jan 19, 2025 8 Share [Cross-posted from LessWrong ] I mostly work on risks from scheming (that is, misaligned, power-seeking AIs that plot against their creators such as by faking alignment). Recently, I (and co-authors) released " Alignment Faking in Large Language Models ", which provides empirical evidence for some components of the scheming threat model. One question that's really important is how likely scheming is. But it's also really important to know how much we expect this
Explore this link on the map →related reading
- How will we update about scheming? — LessWronglesswrong.com
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- AI 2027ai-2027.com
- Frontier Models are Capable of In-context Scheming — AI Alignment Forumalignmentforum.org
- Reducing risk from scheming by studying trained-in scheming behavior — LessWronglesswrong.com
- Catching AIs red-handed — LessWronglesswrong.com
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Training-time schemers vs behavioral schemers — LessWronglesswrong.com