flâneur — a map of the web's best reading

The case for ensuring that powerful AIs are controlled

blog.redwoodresearch.org · 17,210 words · saved by 1 readers

In this post, we argue that AI labs should ensure that powerful AIs are controlled. That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures. We think no fundamental research breakthroughs are required for labs to implement safety measures that meet our standard for AI control for early transformatively useful AIs; we think that meeting our standard would substantially reduce the risks posed by intentional subversion. We recently released a paper in which we explored AI control in a simple programming setting. In that paper, we demonstrated a methodology for evaluating whether any given set of safety measures can control a model. We also introduced several techniques that let us improve the level of safety substantially above baselines without substantially sacrificing performance. In this post, we'll describe AI control in a broade

The case for ensuring that powerful AIs are controlled Labs should make sure that powerful models can't cause unacceptably bad outcomes even if the AIs try to. Buck Shlegeris and Ryan Greenblatt May 07, 2024 20 1 2 Share [Originally posted to LessWrong ] In this post, we argue that AI labs should ensure that powerful AIs are controlled . That is, labs should make sure that the safety measures they apply to their powerful models prevent unacceptably bad outcomes, even if the AIs are misaligned and intentionally try to subvert those safety measures. We think no fundamental research breakthroughs

Explore this link on the map →

saved by

related reading