AI Control: Improving Safety Despite Intentional Subversion
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. As large language models (LLMs) become more powerful and are deployed more autonomously, it will be increasingly important to prevent them from causing harmful outcomes. Researchers have investigated a variety of safety techniques for this purpose, e.g. using models to review the outputs of other models, or red-teaming techniques to surface subtle failure modes. However, rese
AI Control: Improving Safety Despite Intentional Subversion Ryan Greenblatt ∗ Buck Shlegeris Kshitij Sachan Fabien Roger Redwood Research Abstract As large language models (LLMs) become more powerful and are deployed more autonomously, it will be increasingly important to prevent them from causing harmful outcomes. Researchers have investigated a variety of safety techniques for this purpose, e.g. using models to review the outputs of other models, or red-teaming techniques to surface subtle failure modes. However, researchers have not evaluated whether such techniques still ensure safety if t
Explore this link on the map →related reading
- 2312.06942arxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionar5iv.labs.arxiv.org
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- gpt-4.pdfcdn.openai.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Agentic Monitoring for AI Control — LessWronglesswrong.com
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols? — AI Alignment Forumalignmentforum.org