Rohin Shah on what it's really like to run AGI safety at Google DeepMind (and where I disagree with 'doomers') | 80,000 Hours
Building an artifact like Gemini is very, very difficult. The main reason being you have to produce this one thing, this single set of model weights, deployed using a single serving stack. And it has to satisfy so many constraints… There’s probably 100 such things. And it is the case that if you make one change to the process with the intent of making one of these things better — say, safety — it will have random downstream knock-on effects on other constraints that you totally did not anticipate. — Rohin Shah Most people working on AI safety think without a massive effort AI systems will probably end up with goals catastrophically different from humanity’s. Today’s guest, Rohin Shah — head of AGI Safety and Alignment at Google DeepMind, and an AI safety researcher since 2017 — disagrees. “There is no particularly compelling argument that this is the thing that happens by default,” Rohin explains. “There’s a lot of arguments that are suggestive that maybe it could happen, such that you
Rohin Shah on what it's really like to run AGI safety at Google DeepMind (and where I disagree with 'doomers') | 80,000 Hours Search for: On this page: 1 Introduction 1.1 The episode in a nutshell 2 Highlights 3 Articles, books, and other media discussed in the show 4 Transcript 4.1 Who's Rohin Shah? [00:00:00] 4.2 Why Rohin thinks we won't get catastrophic misalignment [00:00:49] 4.3 The limitations of safety and alignment commitments [00:10:38] 4.4 Does Rohin's team have veto power at Google DeepMind? [00:27:36] 4.5 Central banks as a roadmap for regulating AI [00:33:34] 4.6 H
Explore this link on the map →saved by
related reading
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Planning for AGI and beyond | OpenAIopenai.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- AGI safety career advice — EA Forumforum.effectivealtruism.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- AI 2027ai-2027.com
- Project Mario - Colossuscolossus.com
- AI Safety | Arkosevictoriabrook.github.io