Value Learning Needs a Low-Dimensional Bottleneck — LessWrong
Comment by beren - Another line of evidence for the 'values are low-dimensional' is all the emergent misalignment work which tends to find that a.) models have a concept of 'general evil' which goes from writing bad code to giving false medical advice and supporting hitler, and b.) this is controlled often a single or a few directions in the residual stream, which implies an extremely small subspace is behind a model's understanding of morality, and hence (presumably?) the general structure of alignment/morality in the dataset. Emergent misalignment is problematic but it also suggests the possibility of 'emergent alignment' where if a model is trained to be good and aligned in many aspect it may also generalise that far to be aligned in many aspects.
x Value Learning Needs a Low-Dimensional Bottleneck — LessWrong Complexity of value Human Values Inverse Reinforcement Learning AI Frontpage 24 Value Learning Needs a Low-Dimensional Bottleneck by Gunnar_Zarncke 23rd Jan 2026 2 min read 7 24 Epistemic status: Confident in the direction, not confident in the numbers. I have spent a few hours looking into this. Suppose human values were internally coherent, high-dimensional, explicit, and decently stable under reflection. Would alignment be easier or harder? My below calculations show that it would be much harder, if not impossible. I'm going to
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- The Shard Theory of Human Valuesturntrout.com
- Alignment By Default — AI Alignment Forumalignmentforum.org
- What Is The Alignment Problem? — LessWronglesswrong.com
- The Computational Anatomy of Human Values — LessWronglesswrong.com
- The shard theory of human values — LessWronglesswrong.com
- Thomas Larsen's Shortform — LessWronglesswrong.com
- Alignment Is Proven To Be Solvable - by SE Gygesverysane.ai
- Value systematization: how values become coherent (and misaligned) — AI Alignment Forumalignmentforum.org
- Value systematization: how values become coherent (and misaligned) — LessWronglesswrong.com
- Complexity of value — LessWronglesswrong.com