Scheming AIs Will AIs fake alignment during training in order to get power?
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by selecting from this list of supported packages. [online]year This report examines whether advanced AIs that perform well in training will be doing so in order to gain power later – a behavior I call “scheming” (also sometimes called “deceptive alignment”). I conclude that scheming is a disturbingly plausible outcome of using baseline machine learning methods to train goal-directed AIs sophisticated enough t
\DeclareLabeldate [online]year Scheming AIs Will AIs fake alignment during training in order to get power? Joe Carlsmith Open Philanthropy November 2023 Audio version Abstract This report examines whether advanced AIs that perform well in training will be doing so in order to gain power later – a behavior I call “scheming” (also sometimes called “deceptive alignment”). I conclude that scheming is a disturbingly plausible outcome of using baseline machine learning methods to train goal-directed AIs sophisticated enough to scheme (my subjective probability on such an outcome, given these conditi
Explore this link on the map →related reading
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- How will we update about scheming? — LessWronglesswrong.com
- Training-time schemers vs behavioral schemers — LessWronglesswrong.com
- Why AI alignment could be hard with modern deep learningcold-takes.com
- How will we update about scheming? - by Ryan Greenblattredwoodresearch.substack.com
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Frontier Models are Capable of In-context Scheming — AI Alignment Forumalignmentforum.org
- Reducing risk from scheming by studying trained-in scheming behavior — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org