flâneur — a map of the web's best reading

Goal Misgeneralisation: Why Correct Specifications Aren’t Enough For Correct Goals | by DeepMind Safety Research | Medium

deepmindsafetyresearch.medium.com · 2,082 words · saved by 1 readers

By Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton. For more details, check out…

Goal Misgeneralisation: Why Correct Specifications Aren’t Enough For Correct Goals DeepMind Safety Research 9 min read · Oct 7, 2022 -- 1 Listen Share By Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton. For more details, check out our paper . As we build increasingly advanced AI systems, we want to make sure they don’t pursue undesired goals. This is the primary concern of the AI alignment community. Undesired behaviour in an AI agent is often the result of specification gaming —when the AI exploits an incorrectly specified reward. Howev

Explore this link on the map →

related reading