Thoughts on Human Models - LessWrong
Human values and preferences are hard to specify, especially in complex domains. Accordingly, much AGI safety research has focused on approaches to AGI design that refer to human values and preferenc…
x Thoughts on Human Models — LessWrong Research Agendas Mindcrime AI Curated 127 Thoughts on Human Models by Ramana Kumar , Scott Garrabrant 21st Feb 2019 AI Alignment Forum 12 min read 32 127 Ω 51 Human values and preferences are hard to specify, especially in complex domains. Accordingly, much AGI safety research has focused on approaches to AGI design that refer to human values and preferences indirectly , by learning a model that is grounded in expressions of human values (via stated preferences, observed behaviour, approval, etc.) and/or real-world processes that generate expressions of t
Explore this link on the map →related reading
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Planning for AGI and beyond | OpenAIopenai.com
- Commentary on AGI Safety from First Principles — AI Alignment Forumalignmentforum.org
- Where I agree and disagree with Eliezer — LessWronglesswrong.com
- AGI safety career advice — EA Forumforum.effectivealtruism.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- AGI safety from first principles: Introduction — LessWronglesswrong.com
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com
- Rohin Shah on what it's really like to run AGI safety at Google DeepMind (and where I disagree with 'doomers') | 80,000 Hours80000hours.org