How will we update about scheming? — LessWrong
lesswrong.com · 16,383 words · saved by 1 readers
I mostly work on risks from scheming (that is, misaligned, power-seeking AIs that plot against their creators such as by faking alignment). Recently,…
x How will we update about scheming? — LessWrong Redwood Research Deceptive Alignment Outer Alignment AI Curated 2025 Top Fifty: 5 % 177 How will we update about scheming? by ryan_greenblatt 6th Jan 2025 AI Alignment Forum 44 min read 21 177 Ω 86 I mostly work on risks from scheming (that is, misaligned, power-seeking AIs that plot against their creators such as by faking alignment). Recently, I (and co-authors) released " Alignment Faking in Large Language Models ", which provides empirical evidence for some components of the scheming threat model. One question that's really important is how
related reading
- How will we update about scheming?blog.redwoodresearch.org
- How will we update about scheming? - by Ryan Greenblattredwoodresearch.substack.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- Reducing risk from scheming by studying trained-in scheming behavior — LessWronglesswrong.com
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- Frontier Models are Capable of In-context Scheming — AI Alignment Forumalignmentforum.org
- Catching AIs red-handed — LessWronglesswrong.com
- AI 2027ai-2027.com
- Many arguments for AI x-risk are wrong — AI Alignment Forumalignmentforum.org
- Training-time schemers vs behavioral schemers — LessWronglesswrong.com