Goal Misgeneralisation: Why Correct Specifications Aren’t Enough For Correct Goals | by DeepMind Safety Research | Medium
deepmindsafetyresearch.medium.com · 2,082 words · saved by 1 readers
By Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton. For more details, check out…
Goal Misgeneralisation: Why Correct Specifications Aren’t Enough For Correct Goals DeepMind Safety Research 9 min read · Oct 7, 2022 -- 1 Listen Share By Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton. For more details, check out our paper . As we build increasingly advanced AI systems, we want to make sure they don’t pursue undesired goals. This is the primary concern of the AI alignment community. Undesired behaviour in an AI agent is often the result of specification gaming —when the AI exploits an incorrectly specified reward. Howev
related reading
- Specification gaming: the flip side of AI ingenuity — Google DeepMinddeepmind.google
- Deep Deceptiveness — LessWronglesswrong.com
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org
- Specification gaming: the flip side of AI ingenuity — Google DeepMinddeepmind.google
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Model Spec Midtraining: Improving How Alignment Training Generalizesalignment.anthropic.com
- Why AI alignment could be hard with modern deep learningcold-takes.com
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com
- When does training a model change its goals?blog.redwoodresearch.org
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org
- Specification gaming examples in AI — LessWronglesswrong.com