New report: "Scheming AIs: Will AIs fake alignment during training in order to get power?" - Joe Carlsmith
joecarlsmith.com · 9,546 words · saved by 3 readers
My report examining the probability of a behavior often called "deceptive alignment."
Podcast version here, or search “Joe Carlsmith Audio” on your podcast app I’ve written a report about whether advanced AIs will fake alignment during training in order to get power later – a behavior I call “scheming” (also sometimes called “deceptive alignment”). The report is available on arXiv here. There’s also an audio version here, and I’ve included the introductory section below. This section includes a full summary of the report, which covers most of the main points and technical terminology. I’m hoping that the summary will provide much of the context necessary to understand…
saved by
related reading
- How will we update about scheming?blog.redwoodresearch.org
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- Deep Deceptiveness — LessWronglesswrong.com
- Many arguments for AI x-risk are wrong — AI Alignment Forumalignmentforum.org
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- How will we update about scheming? — LessWronglesswrong.com
- Why AI alignment could be hard with modern deep learningcold-takes.com