[2607.18506] AI Value Alignment for Evolving Social Norms
Abstract:AI alignment is essential for the safe deployment of advanced AI systems. Given that values and preferences change over time, culture, social roles, and context, we need to develop a better understanding of the possible long-term consequences of AI alignment, in particular considering the likely ubiquitous future use of personalized AI assistants. We introduce a flexible and extensible mathematical modelling framework, rooted in social physics, aimed at answering macro-level questions regarding the evolving social norms in human populations under the assumption of frequent AI use. Our analysis is part-analytical, and part-simulation, enabling us to characterize the long-term dynamical consequences under a diverse set of starting assumptions. We highlight the risk of value lock-in, and normative mode collapse, prominently featured in non-adaptive alignment formulations. Beyond alignment, we advocate for the wider adoption of these kinds of social physics models as an epistemic bridge: enabling rapid, rigorous, and quantitatively-grounded hypothesis testing for sociotechnical foresight in general AI futures, and acting as a tractable precursor to more computationally expensive large-scale agentic evaluations.
View PDF HTML (experimental) Abstract:AI alignment is essential for the safe deployment of advanced AI systems. Given that values and preferences change over time, culture, social roles, and context, we need to develop a better understanding of the possible long-term consequences of AI alignment, in particular considering the likely ubiquitous future use of personalized AI assistants. We introduce a flexible and extensible mathematical modelling framework, rooted in social physics, aimed at answering macro-level questions regarding the evolving social norms in human populations under the…
saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Full-Stack Alignment and Thick Models of Valuefull-stack-alignment.ai
- The Artificiality of Alignmentjoinreboot.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- [2605.10310] Positive Alignment: Artificial Intelligence for Human Flourishingarxiv.org
- Gradual Disempowermentgradual-disempowerment.ai
- Societal and technological progress as sewing an ever-growing, ever-changing, patchy, and polychrome quilt — LessWronglesswrong.com
- What are we building for? Artificial intelligence, alignment, and a future worth wantingmag.re-alignment.com
- A Long Sequence of Small, Correct Decisionscalebwatney.substack.com
- Another (outer) alignment failure story — AI Alignment Forumalignmentforum.org
- Value Alignment Is a Pseudo Conceptzilanqian.substack.com