Notes on handling non-concentrated failures with AI control: high level methods and different regimes
blog.redwoodresearch.org · 5,088 words · saved by 1 readers
What are the methods and issues when failures occur diffusely over many actions?
In this post, I'll try to explain my current understanding of the high level methods for handling non-concentrated failures with control. I'll discuss the regimes produced by different methods and the failure modes of these different regimes. Non-concentrated failures are issues that arise from the AI doing a large number of problematic actions over a long period. (See here and here for more discussion of concentrated vs non-concentrated failures.) A key aspect of non-concentrated failures is that they can in principle be handled by applying asynchronous online training with a good enough…
saved by
related reading
- Auditing failures vs concentrated failures — AI Alignment Forumalignmentforum.org
- Catching AIs red-handedblog.redwoodresearch.org
- Notes on handling non-concentrated failures with AI control: high level methods and different regimes — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- What failure looks like — LessWronglesswrong.com
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Reading Listblog.redwoodresearch.org