Auditing failures vs concentrated failures — AI Alignment Forum
TL;DR: If you want to notice when powerful AIs are behaving badly, a reasonable strategy is to carefully audit a small fraction of AI actions. This kind of strategy has two kinds of failures which can be tackled somewhat independently: auditing failures, where the audits fail to recognize some bad behavior, and concentrated failures, where the AI causes an irreversible catastrophe in so few actions or such a short time frame that auditing didn’t have a chance to review the bad actions. [The point made in this post is somewhat narrow and technical, but we often want to reference this point.] Many oversight techniques are so expensive that using them on every AI action would destroy most of the value of using AIs in the first place: you might want to have humans extensively review all AI-generated code line by line, but this reduces AI usefulness so much that you would often be better off having humans write the code in the first place. A typical way to solve this problem is to just audi
x Auditing failures vs concentrated failures — AI Alignment Forum AI Frontpage 25 Auditing failures vs concentrated failures by ryan_greenblatt , Fabien Roger 11th Dec 2023 9 min read 2 25 TL;DR : If you want to notice when powerful AIs are behaving badly, a reasonable strategy is to carefully audit a small fraction of AI actions. This kind of strategy has two kinds of failures which can be tackled somewhat independently: auditing failures , where the audits fail to recognize some bad behavior, and concentrated failures , where the AI causes an irreversible catastrophe in so few actions or suc
Explore this link on the map →saved by
related reading
- Catching AIs red-handedblog.redwoodresearch.org
- AI catastrophes and rogue deployments - by Buck Shlegerisblog.redwoodresearch.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Notes on handling non-concentrated failures with AI control: high level methods and different regimes — LessWronglesswrong.com
- What failure looks like — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Win/continue/lose scenarios and execute/replace/audit protocols — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Another (outer) alignment failure story — AI Alignment Forumalignmentforum.org
- Building and evaluating alignment auditing agents — AI Alignment Forumalignmentforum.org