flâneur

How will we update about scheming? — LessWrong

lesswrong.com · 16,383 words · saved by 1 readers

I mostly work on risks from scheming (that is, misaligned, power-seeking AIs that plot against their creators such as by faking alignment). Recently,…

x How will we update about scheming? — LessWrong Redwood Research Deceptive Alignment Outer Alignment AI Curated 2025 Top Fifty: 5 % 177 How will we update about scheming? by ryan_greenblatt 6th Jan 2025 AI Alignment Forum 44 min read 21 177 Ω 86 I mostly work on risks from scheming (that is, misaligned, power-seeking AIs that plot against their creators such as by faking alignment). Recently, I (and co-authors) released " Alignment Faking in Large Language Models ", which provides empirical evidence for some components of the scheming threat model. One question that's really important is how

related reading