✳flâneur — a map of the web's best reading
AI Alignment Podcast: Cooperative Inverse Reinforcement Learning with Dylan Hadfield-Menell (Beneficial AGI 2019) - Future of Life Institute
futureoflife.org · 10,283 words · saved by 1 readers
In this episode of the AI alignment series, Lucas interviews Dylan Hadfield-Menell on cooperative inverse reinforcement learning.
AI Alignment Podcast: Cooperative Inverse Reinforcement Learning with Dylan Hadfield-Menell (Beneficial AGI 2019) - Future of Life Institute Skip to content All Podcast Episodes AI Alignment Podcast: Cooperative Inverse Reinforcement Learning with Dylan Hadfield-Menell (Beneficial AGI 2019) Published 17 January, 2019 What motivates cooperative inverse reinforcement learning? What can we gain from recontextualizing our safety efforts from the CIRL point of view? What possible role can pre-AGI systems play in amplifying normative processes? Cooperative Inverse Reinforcement Learning with Dylan H
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- [2605.10310] Positive Alignment: Artificial Intelligence for Human Flourishingarxiv.org
- (My understanding of) What Everyone in Technical Alignment is Doing and Why — LessWronglesswrong.com
- [AN #70]: Agents that help humans who are still learning about their own preferences — LessWronglesswrong.com
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org
- Shah and Yudkowsky on alignment failures — LessWronglesswrong.com
- Research Publications – Center for Human-Compatible Artificial Intelligencehumancompatible.ai
- LLM Alignment, ethical and mathematical realism, and the most important actions in davidad's understanding — LessWronglesswrong.com