Thoughts on the conservative assumptions in AI control
Work that I’ve done on techniques for mitigating risk from misaligned AI often makes a number of conservative assumptions about the capabilities of the AIs we’re trying to control. (E.g. the original AI control paper, Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats, How to prevent collusion when using untrusted models to monitor each other.) For example: The AIs are consistently trying to subvert safety measures. They’re very good at strategizing about how to subvert them, and can do this strategizing totally illegibly (e.g. in a single forward pass). They can follow a strategy that’s coordinated between all their instances. They know exactly what safety measures are in use. They know a lot about the distribution of inputs. They’re able to pick Schelling points perfectly. They’re able to follow stochastic policies with excellent calibration (e.g. they’re able to attack with probability 0.01%). But I don’t think situations where all these assumptions hold contain the m
Work that I’ve done on techniques for mitigating risk from misaligned AI often makes a number of conservative assumptions about the capabilities of the AIs we’re trying to control. (E.g. the original AI control paper, Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats, How to prevent collusion when using untrusted models to monitor each other.) For example: The AIs are consistently trying to subvert safety measures. They’re very good at strategizing about how to subvert them, and can do this strategizing totally illegibly (e.g. in a single forward pass). They can follow a…
related reading
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Reading Listblog.redwoodresearch.org
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Introduction to AI Control - by Sarah - BlueDot Impactblog.bluedot.org
- The Case Against AI Control Research — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — AI Alignment Forumalignmentforum.org
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Adolescence of Technologydarioamodei.com