flâneur — a map of the web's best reading

Value Learning Needs a Low-Dimensional Bottleneck — LessWrong

lesswrong.com · 1,649 words · saved by 1 readers

Comment by beren - Another line of evidence for the 'values are low-dimensional' is all the emergent misalignment work which tends to find that a.) models have a concept of 'general evil' which goes from writing bad code to giving false medical advice and supporting hitler, and b.) this is controlled often a single or a few directions in the residual stream, which implies an extremely small subspace is behind a model's understanding of morality, and hence (presumably?) the general structure of alignment/morality in the dataset. Emergent misalignment is problematic but it also suggests the possibility of 'emergent alignment' where if a model is trained to be good and aligned in many aspect it may also generalise that far to be aligned in many aspects.

x Value Learning Needs a Low-Dimensional Bottleneck — LessWrong Complexity of value Human Values Inverse Reinforcement Learning AI Frontpage 24 Value Learning Needs a Low-Dimensional Bottleneck by Gunnar_Zarncke 23rd Jan 2026 2 min read 7 24 Epistemic status: Confident in the direction, not confident in the numbers. I have spent a few hours looking into this. Suppose human values were internally coherent, high-dimensional, explicit, and decently stable under reflection. Would alignment be easier or harder? My below calculations show that it would be much harder, if not impossible. I'm going to

Explore this link on the map →

related reading