✳flâneur — a map of the web's best reading
Thoughts on the conservative assumptions in AI control
blog.redwoodresearch.org · 3,819 words · saved by 2 readers
Why are we so friendly to the red team?
Thoughts on the conservative assumptions in AI control Why are we so friendly to the red team? Buck Shlegeris Jan 17, 2025 5 Share Work that I’ve done on techniques for mitigating risk from misaligned AI often makes a number of conservative assumptions about the capabilities of the AIs we’re trying to control. (E.g. the original AI control paper , Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats , How to prevent collusion when using untrusted models to monitor each other .) For example: The AIs are consistently trying to subvert safety measures. They’re very good at strategizi
Explore this link on the map →saved by
related reading
- How can we solve diffuse threats like research sabotage with AI control?blog.redwoodresearch.org
- The Case Against AI Control Research — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org
- The case for ensuring that powerful AIs are controlled — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org