How can we solve diffuse threats like research sabotage with AI control? — LessWrong
Misaligned AIs might engage in research sabotage: making safety research go poorly by doing things like withholding their best ideas or putting subtl…
x How can we solve diffuse threats like research sabotage with AI control? — LessWrong AI Control AI Frontpage 54 How can we solve diffuse threats like research sabotage with AI control? by Vivek Hebbar 30th Apr 2025 AI Alignment Forum 10 min read 1 54 Ω 34 Misaligned AIs might engage in research sabotage: making safety research go poorly by doing things like withholding their best ideas or putting subtle bugs in experiments. To mitigate this risk with AI control, we need very different techniques than those in the original control paper or Ctrl-Z , because research sabotage differs from those
Explore this link on the map →saved by
related reading
- How can we solve diffuse threats like research sabotage with AI control?blog.redwoodresearch.org
- How can we solve diffuse threats like research sabotage with AI control?redwoodresearch.substack.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- The Case Against AI Control Research — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- gdm-ai-control-roadmap.pdfstorage.googleapis.com
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- Fields that I reference when thinking about AI takeover prevention — LessWronglesswrong.com
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org