Neel Nanda on the race to read AI minds (part 1) - 80,000 Hours
It’s kind of wild that we can do chain-of-thought monitoring. … We’ve been scared of these terrifying black box systems that will do inscrutable things — and they just think in English? And there are actual reasons to think that looking at those thoughts will be informative rather than bullshit? What? This is so convenient, yet it is also very fragile. And it’s very important that we don’t do things that would make it no longer monitorable. — Neel Nanda We don’t know how AIs think or why they do what they do. Or at least, we don’t know much. That fact is only becoming more troubling as AIs grow more capable and appear on track to wield enormous cultural influence, directly advise on major government decisions, and even operate military equipment autonomously. We simply can’t tell what models, if any, should be trusted with such authority. Neel Nanda of Google DeepMind is one of the founding figures of the field of machine learning trying to fix this situation — mechanistic interpretabi
Neel Nanda on the race to read AI minds (part 1) | 80,000 Hours Search for: On this page: 1 Introduction 1.1 The episode in a nutshell 2 Highlights 3 Articles, books, and other media discussed in the show 4 Transcript 4.1 Cold open [00:00:00] 4.2 Who's Neel Nanda? [00:01:04] 4.3 How would mechanistic interpretability help with AGI [00:02:01] 4.4 What's mech interp? [00:05:12] 4.5 How Neel changed his take on mech interp [00:09:50] 4.6 Top successes in interpretability [00:16:00] 4.7 Probes can cheaply detect harmful intentions in AIs [00:20:13] 4.8 In some ways we understand AIs better than hu
Explore this link on the map →related reading
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- On Optimism for Interpretabilitygoodfire.ai
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com
- How AI Is Learning to Think in Secret — LessWronglesswrong.com
- Natural Language Autoencoders \ Anthropicanthropic.com