2404.10636
arxiv.org · 8,027 words · saved by 1 readers
N/A
What are human values, and how do we align AI to them? Oliver Klingefjord Ryan Lowe∗ Joe Edelman arXiv:2404.10636v2 [cs.CY] 17 Apr 2024 Meaning Alignment Institute Abstract There is an emerging consensus…
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- The Artificiality of Alignmentjoinreboot.org
- Alignment Is Proven To Be Solvable - by SE Gygesverysane.ai
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- Teaching Claude Whyalignment.anthropic.com
- Full-Stack Alignment and Thick Models of Valuefull-stack-alignment.ai
- Value Alignment Is a Pseudo Conceptzilanqian.substack.com
- [2001.09768] Artificial Intelligence, Values and Alignmentarxiv.org
- [2008.02275] Aligning AI With Shared Human Valuesarxiv.org
- [2605.10310] Positive Alignment: Artificial Intelligence for Human Flourishingarxiv.org
- Teaching Claude why \ Anthropicanthropic.com
- [2112.00861] A General Language Assistant as a Laboratory for Alignmentarxiv.org