Thoughts on the conservative assumptions in AI control
blog.redwoodresearch.org · 3,819 words · saved by 3 readers
Why are we so friendly to the red team?
Thoughts on the conservative assumptions in AI control Why are we so friendly to the red team? Buck Shlegeris Jan 17, 2025 5 Share Work that I’ve done on techniques for mitigating risk from misaligned AI often makes a number of conservative assumptions about the capabilities of the AIs we’re trying to control. (E.g. the original AI control paper , Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats , How to prevent collusion when using untrusted models to monitor each other .) For example: The AIs are consistently trying to subvert safety measures. They’re very good at strategizi
saved by
related reading
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- The Case Against AI Control Research — LessWronglesswrong.com
- Thoughts on the conservative assumptions in AI controlredwoodresearch.substack.com
- Reading Listblog.redwoodresearch.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- How can we solve diffuse threats like research sabotage with AI control?blog.redwoodresearch.org
- Catching AIs red-handedblog.redwoodresearch.org
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- Why it's hard to make settings for high-stakes control researchblog.redwoodresearch.org
- AI 2040: Plan Aai-2040.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com