Towards Bidirectional Human-AI Alignment
arxiv.org · 5,918 words · saved by 1 readers
N/A
Position: Towards Bidirectional Human-AI Alignment Hua Shen1A ∗ Tiffany Knearem2B Reshmi Ghosh3B 4C Kenan Alkiek Kundan Krishna5C Yachuan Liu4C Ziqiao Ma4C Savvas Petridis9C 7C Yi-Hao Peng Li Qiwei4C Sushrita Rakshit4C Chenglei Si8C Yutong Xie4C…
related reading
- The Artificiality of Alignmentjoinreboot.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- [2605.10310] Positive Alignment: Artificial Intelligence for Human Flourishingarxiv.org
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Value Alignment Is a Pseudo Conceptzilanqian.substack.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Explaining AI Alignment as an NLPer and Why I am Working on Itruiqizhong.substack.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- Full-Stack Alignment and Thick Models of Valuefull-stack-alignment.ai