flâneur — a map of the web's best reading

Auditing failures vs concentrated failures — AI Alignment Forum

alignmentforum.org · 2,565 words · saved by 1 readers

TL;DR: If you want to notice when powerful AIs are behaving badly, a reasonable strategy is to carefully audit a small fraction of AI actions. This kind of strategy has two kinds of failures which can be tackled somewhat independently: auditing failures, where the audits fail to recognize some bad behavior, and concentrated failures, where the AI causes an irreversible catastrophe in so few actions or such a short time frame that auditing didn’t have a chance to review the bad actions. [The point made in this post is somewhat narrow and technical, but we often want to reference this point.] Many oversight techniques are so expensive that using them on every AI action would destroy most of the value of using AIs in the first place: you might want to have humans extensively review all AI-generated code line by line, but this reduces AI usefulness so much that you would often be better off having humans write the code in the first place. A typical way to solve this problem is to just audi

x Auditing failures vs concentrated failures — AI Alignment Forum AI Frontpage 25 Auditing failures vs concentrated failures by ryan_greenblatt , Fabien Roger 11th Dec 2023 9 min read 2 25 TL;DR : If you want to notice when powerful AIs are behaving badly, a reasonable strategy is to carefully audit a small fraction of AI actions. This kind of strategy has two kinds of failures which can be tackled somewhat independently: auditing failures , where the audits fail to recognize some bad behavior, and concentrated failures , where the AI causes an irreversible catastrophe in so few actions or suc

Explore this link on the map →

saved by

related reading