flâneur — a map of the web's best reading

What is AI alignment? – BlueDot Impact

aisafetyfundamentals.com · saved by 1 readers

There are many competing definitions for terms in AI safety, especially for ‘alignment’. This short article explains how we define these words on our AI Alignment course, as well as how alignment contributes to AI safety alongside other research areas. AI safety is concerned with reducing AI risks, ultimately decreasing the expected harm from AI systems. We mean this in a very broad sense, and correspondingly many subfields contribute to this. One way these can be divided up is: Alignment: making AI systems try to do what their creators intend them to do (some people call this intent alignment). Examples of misalignment: image generators create images that exacerbate stereotypes or are unrealistically diverse, medical classifiers look for rulers rather than medically relevant features, and chatbots tell people what they want to hear rather than the truth. In the future, we might delegate more power to extremely capable systems - if these systems are not doing what we intend them to do,

There are many competing definitions for terms in AI safety, especially for ‘alignment’. This short article explains how we define these words on our AI Alignment course, as well as how alignment contributes to AI safety alongside other research areas. AI safety is concerned with reducing AI risks, ultimately decreasing the expected harm from AI systems. We mean this in a very broad sense, and correspondingly many subfields contribute to this. One way these can be divided up is: Alignment: making AI systems try to do what their creators intend them to do (some people call this intent alignment

Explore this link on the map →