What is AI alignment? – BlueDot Impact
There are many competing definitions for terms in AI safety, especially for ‘alignment’. This short article explains how we define these words on our AI Alignment course, as well as how alignment contributes to AI safety alongside other research areas. AI safety is concerned with reducing AI risks, ultimately decreasing the expected harm from AI systems. We mean this in a very broad sense, and correspondingly many subfields contribute to this. One way these can be divided up is: Alignment: making AI systems try to do what their creators intend them to do (some people call this intent alignment). Examples of misalignment: image generators create images that exacerbate stereotypes or are unrealistically diverse, medical classifiers look for rulers rather than medically relevant features, and chatbots tell people what they want to hear rather than the truth. In the future, we might delegate more power to extremely capable systems - if these systems are not doing what we intend them to do,
This article explains key concepts that come up in the context of AI alignment. These terms are only attempts at gesturing at the underlying ideas, and the ideas are what is important. There is no strict consensus on which name should correspond to which idea, and different people use the terms differently.1 This article explains how we use these words on our AI Alignment course, and how alignment research contributes to AI safety. AI is likely to have an unprecedented impact on the world — possibly comparable to the industrial revolution or even larger. Terms to describe these systems…
related reading
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org
- What is AI alignment? - by Adam Jones - BlueDot Impactaisafetyfundamentals.com
- What is AI alignment? - by Adam Jones - BlueDot Impactbluedot.org
- The Artificiality of Alignmentjoinreboot.org
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- AI safety - Wikipediaen.wikipedia.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Where I agree and disagree with Eliezer — LessWronglesswrong.com
- [2605.10310] Positive Alignment: Artificial Intelligence for Human Flourishingarxiv.org
- AI Safety for Fleshy Humans: a whirlwind touraisafety.dance