Introduction to AI Control | BlueDot Impact
AI Control is a research agenda that aims to prevent misaligned AI systems from causing harm. It is different from AI alignment, which aims to ensure that systems act in the best interests of their users. Put simply, aligned AIs do not want to harm humans, whereas controlled AIs can't harm humans, even if they want to. There are a few reasons why control-style research could be useful for AI safety. Some believe that AI control might be easier than AI alignment, at least in the short term. A core challenge in aligning AI systems is the possibility of deception. Models may appear to be acting in our interests during training and evaluation, but they could actually be harbouring ulterior motives. Because we don't understand much of what is happening inside neural networks, deception could be hard to detect. This could lead us to deploy models that are actually scheming against us. But if our aim is merely to control, rather than to align, a model, we don't need to know whether a model wa
Blog Introduction to AI Control Sarah Apr 26, 2025 35 2 Share AI Control is a research agenda that aims to prevent misaligned AI systems from causing harm. It is different from AI alignment , which aims to ensure that systems act in the best interests of their users. Put simply, aligned AIs do not want to harm humans, whereas controlled AIs can’t harm humans, even if they want to. Why might AI Control be useful? There are a few reasons why control-style research could be useful for AI safety. AI control might be easier than AI alignment Some believe that AI control might be easier than AI alig
related reading
- Introduction to AI Control - by Sarah - BlueDot Impactblog.bluedot.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- The Case Against AI Control Research — LessWronglesswrong.com
- Reading Listblog.redwoodresearch.org
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- I think alignment work is more promising than control work — LessWronglesswrong.com
- How useful is AI control? @ Trackstracks.xlabtracks.workers.dev
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- An overview of areas of control workblog.redwoodresearch.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — AI Alignment Forumalignmentforum.org