Against cultural alignment - by Harry Law
Whenever someone mentions AI alignment you can bet that someone else isn’t too far away from asking ‘alignment to what?’ with a certain degree of satisfaction. I’m thinking about coining a new law of internet discourse to describe this phenomenon. Something like Godwin’s law but for posts about AI. For those scratching their heads, it’s funny and frustrating in equal measure because it muddles two types of alignment. There are lots of different ways to describe these groups, but for our purposes we can think of them as technical alignment and value alignment. The former deals with ‘getting AI to do what you want’. This is the problem that labs try to solve with gigantic sticking plasters like reinforcement learning from human feedback (RLHF), where the model is steered to interpret instructions, avoid jailbreaks, and generally avoid the spectacle of crashing out. Our second species of alignment asks whether an AI’s actions are ethically appropriate, and wants to know whose values they
Essays Against cultural alignment AI with local values sounds great. So what's the problem? Harry Law Jul 01, 2025 29 11 5 Share Fresco of constellations in Palazzo Farnese by Giovanni de' Vecchi 1574 Whenever someone mentions AI alignment you can bet that someone else isn’t too far away from asking ‘ alignment to what ?’ with a certain degree of satisfaction. I’m thinking about coining a new law of internet discourse to describe this phenomenon. Something like Godwin’s law but for posts about AI. For those scratching their heads, it’s funny and frustrating in equal measure because it muddles
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Gradual Disempowermentgradual-disempowerment.ai
- The case against AI alignment — LessWronglesswrong.com
- [2605.10310] Positive Alignment: Artificial Intelligence for Human Flourishingarxiv.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- What’s Your P(WEIRD)? — LessWronglesswrong.com
- The Problem With the Word ‘Alignment’ — AI Alignment Forumalignmentforum.org
- Language Log >> Interpersonal and socio-cultural alignmentlanguagelog.ldc.upenn.edu
- Why AI alignment could be hard with modern deep learningcold-takes.com
- The self-unalignment problem — LessWronglesswrong.com
- AI Alignment Cannot Be Top-Down | AI Frontiersai-frontiers.org