[2312.06942] AI Control: Improving Safety Despite Intentional Subversion
As large language models (LLMs) become more powerful and are deployed more autonomously, it will be increasingly important to prevent them from causing harmful outcomes. Researchers have investigated a variety of safet…
AI Control: Improving Safety Despite Intentional Subversion Ryan Greenblatt ∗ Buck Shlegeris Kshitij Sachan Fabien Roger Redwood Research Abstract As large language models (LLMs) become more powerful and are deployed more autonomously, it will be increasingly important to prevent them from causing harmful outcomes. Researchers have investigated a variety of safety techniques for this purpose, e.g. using models to review the outputs of other models, or red-teaming techniques to surface subtle failure modes. However, researchers have not evaluated whether such techniques still ensure safety if t
related reading
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- 2312.06942arxiv.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- gpt-4.pdfcdn.openai.com
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Reading Listblog.redwoodresearch.org
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org