How will we update about scheming? - by Ryan Greenblatt
blog.redwoodresearch.org · 9,647 words · saved by 1 readers
A quantitative description of how I expect to change my mind.
[Cross-posted from LessWrong] I mostly work on risks from scheming (that is, misaligned, power-seeking AIs that plot against their creators such as by faking alignment). Recently, I (and co-authors) released "Alignment Faking in Large Language Models", which provides empirical evidence for some components of the scheming threat model. One question that's really important is how likely scheming is. But it's also really important to know how much we expect this uncertainty to be resolved by various key points in the future. I think it's about 25% likely that the first AIs capable of…
saved by
related reading
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- How will we update about scheming? — LessWronglesswrong.com
- How will we update about scheming? - by Ryan Greenblattredwoodresearch.substack.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- Frontier Models are Capable of In-context Scheming — AI Alignment Forumalignmentforum.org
- AI 2027ai-2027.com
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- Many arguments for AI x-risk are wrong — AI Alignment Forumalignmentforum.org
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- Reducing risk from scheming by studying trained-in scheming behavior — LessWronglesswrong.com