What is AI alignment? – BlueDot Impact
There are many competing definitions for terms in AI safety, especially for ‘alignment’. This short article explains how we define these words on our AI Alignment course, as well as how alignment contributes to AI safety alongside other research areas. AI safety is concerned with reducing AI risks, ultimately decreasing the expected harm from AI systems. We mean this in a very broad sense, and correspondingly many subfields contribute to this. One way these can be divided up is: Alignment: making AI systems try to do what their creators intend them to do (some people call this intent alignment). Examples of misalignment: image generators create images that exacerbate stereotypes or are unrealistically diverse, medical classifiers look for rulers rather than medically relevant features, and chatbots tell people what they want to hear rather than the truth. In the future, we might delegate more power to extremely capable systems - if these systems are not doing what we intend them to do,
There are many competing definitions for terms in AI safety, especially for ‘alignment’. This short article explains how we define these words on our AI Alignment course, as well as how alignment contributes to AI safety alongside other research areas. AI safety is concerned with reducing AI risks, ultimately decreasing the expected harm from AI systems. We mean this in a very broad sense, and correspondingly many subfields contribute to this. One way these can be divided up is: Alignment: making AI systems try to do what their creators intend them to do (some people call this intent alignment
Explore this link on the map →