Model Plurality
Current research in “plural alignment” concentrates on making AI models amenable to diverse human values. But plurality is not simply a safeguard against bias or an engine of efficiency: it's a key ingredient for intelligence itself. Or copy link Efforts to align AI to what is safe or desirable consistently run into the problem of diverse perspectives on what exactly constitutes safe or desirable. Rather than facing the decision of ranking one worldview over another, a growing number of machine learning researchers have instead called for building pluralistic AI, hoping to cram a single model with a set of interchangeable value systems spanning the entire repertoire of cultural moral primitives. These efforts are focused on achieving values plurality: making large language models (LLMs) like ChatGPT amenable to a spread of human values rather than biased towards a single, politicized perspective. “Pluralistic alignment” research defines several possible forms of values plurality in mod
Or copy link Copy By Christina Lu Efforts to align AI to what is safe or desirable consistently run into the problem of diverse perspectives on what exactly constitutes safe or desirable . Rather than facing the decision of ranking one worldview over another, a growing number of machine learning researchers have instead called for building pluralistic AI, hoping to cram a single model with a set of interchangeable value systems spanning the entire repertoire of cultural moral primitives. These efforts are focused on achieving values plurality : making large language models (LLMs) like ChatGPT
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- The Future Worth Building Is Human - Thinking Machines Labthinkingmachines.ai
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- [2605.10310] Positive Alignment: Artificial Intelligence for Human Flourishingarxiv.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- AI #24: Week of the Podcast — LessWronglesswrong.com
- Questions about the Future of AI - by Dwarkesh Pateldwarkesh.com
- Why I’m optimistic about our alignment approachaligned.substack.com
- Why I’m optimistic about our alignment approachaligned.substack.com