✳flâneur — a map of the web's best reading
Goal Misgeneralisation: Why Correct Specifications Aren’t Enough For Correct Goals | by DeepMind Safety Research | Medium
deepmindsafetyresearch.medium.com · 2,082 words · saved by 1 readers
By Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton. For more details, check out…
Goal Misgeneralisation: Why Correct Specifications Aren’t Enough For Correct Goals DeepMind Safety Research 9 min read · Oct 7, 2022 -- 1 Listen Share By Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton. For more details, check out our paper . As we build increasingly advanced AI systems, we want to make sure they don’t pursue undesired goals. This is the primary concern of the AI alignment community. Undesired behaviour in an AI agent is often the result of specification gaming —when the AI exploits an incorrectly specified reward. Howev
Explore this link on the map →related reading
- Specification gaming: the flip side of AI ingenuity — Google DeepMinddeepmind.google
- Specification gaming: the flip side of AI ingenuity — Google DeepMinddeepmind.google
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Why AI alignment could be hard with modern deep learningcold-takes.com
- Specification gaming examples in AI — LessWronglesswrong.com
- AI Goals Forecast — AI 2027ai-2027.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Commentary on AGI Safety from First Principles — AI Alignment Forumalignmentforum.org
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com
- Specification gaming: the flip side of AI ingenuity — Google DeepMinddeepmind.com
- Rohin Shah on what it's really like to run AGI safety at Google DeepMind (and where I disagree with 'doomers') | 80,000 Hours80000hours.org