[2312.06942] AI Control: Improving Safety Despite Intentional Subversion
As large language models (LLMs) become more powerful and are deployed more autonomously, it will be increasingly important to prevent them from causing harmful outcomes. Researchers have investigated a variety of safet…
AI Control: Improving Safety Despite Intentional Subversion Ryan Greenblatt ∗ Buck Shlegeris Kshitij Sachan Fabien Roger Redwood Research Abstract As large language models (LLMs) become more powerful and are deployed more autonomously, it will be increasingly important to prevent them from causing harmful outcomes. Researchers have investigated a variety of safety techniques for this purpose, e.g. using models to review the outputs of other models, or red-teaming techniques to surface subtle failure modes. However, researchers have not evaluated whether such techniques still ensure safety if t
Explore this link on the map →related reading
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- 2312.06942arxiv.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- gpt-4.pdfcdn.openai.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org