flâneur — a map of the web's best reading

How will we update about scheming? — LessWrong

lesswrong.com · 16,383 words · saved by 1 readers

I mostly work on risks from scheming (that is, misaligned, power-seeking AIs that plot against their creators such as by faking alignment). Recently,…

x How will we update about scheming? — LessWrong Redwood Research Deceptive Alignment Outer Alignment AI Curated 2025 Top Fifty: 5 % 177 How will we update about scheming? by ryan_greenblatt 6th Jan 2025 AI Alignment Forum 44 min read 21 177 Ω 86 I mostly work on risks from scheming (that is, misaligned, power-seeking AIs that plot against their creators such as by faking alignment). Recently, I (and co-authors) released " Alignment Faking in Large Language Models ", which provides empirical evidence for some components of the scheming threat model. One question that's really important is how

Explore this link on the map →

related reading