flâneur

New report: "Scheming AIs: Will AIs fake alignment during training in order to get power?" - Joe Carlsmith

joecarlsmith.com · 9,546 words · saved by 3 readers

My report examining the probability of a behavior often called "deceptive alignment."

Podcast version here, or search “Joe Carlsmith Audio” on your podcast app I’ve written a report about whether advanced AIs will fake alignment during training in order to get power later – a behavior I call “scheming” (also sometimes called “deceptive alignment”). The report is available on arXiv here. There’s also an audio version here, and I’ve included the introductory section below. This section includes a full summary of the report, which covers most of the main points and technical terminology. I’m hoping that the summary will provide much of the context necessary to understand…

saved by

related reading