flâneur — a map of the web's best reading

[2312.06942] AI Control: Improving Safety Despite Intentional Subversion

ar5iv.labs.arxiv.org · 20,846 words · saved by 1 readers

As large language models (LLMs) become more powerful and are deployed more autonomously, it will be increasingly important to prevent them from causing harmful outcomes. Researchers have investigated a variety of safet…

AI Control: Improving Safety Despite Intentional Subversion Ryan Greenblatt ∗ Buck Shlegeris Kshitij Sachan Fabien Roger Redwood Research Abstract As large language models (LLMs) become more powerful and are deployed more autonomously, it will be increasingly important to prevent them from causing harmful outcomes. Researchers have investigated a variety of safety techniques for this purpose, e.g. using models to review the outputs of other models, or red-teaming techniques to surface subtle failure modes. However, researchers have not evaluated whether such techniques still ensure safety if t

Explore this link on the map →

related reading