flâneur

How will we update about scheming? - by Ryan Greenblatt

blog.redwoodresearch.org · 9,647 words · saved by 1 readers

A quantitative description of how I expect to change my mind.

[Cross-posted from LessWrong] I mostly work on risks from scheming (that is, misaligned, power-seeking AIs that plot against their creators such as by faking alignment). Recently, I (and co-authors) released "Alignment Faking in Large Language Models", which provides empirical evidence for some components of the scheming threat model. One question that's really important is how likely scheming is. But it's also really important to know how much we expect this uncertainty to be resolved by various key points in the future. I think it's about 25% likely that the first AIs capable of…

saved by

related reading