flâneur — a map of the web's best reading

Neel Nanda on the race to read AI minds (part 1) - 80,000 Hours

80000hours.org · 40,515 words · saved by 1 readers

It’s kind of wild that we can do chain-of-thought monitoring. … We’ve been scared of these terrifying black box systems that will do inscrutable things — and they just think in English? And there are actual reasons to think that looking at those thoughts will be informative rather than bullshit? What? This is so convenient, yet it is also very fragile. And it’s very important that we don’t do things that would make it no longer monitorable. — Neel Nanda We don’t know how AIs think or why they do what they do. Or at least, we don’t know much. That fact is only becoming more troubling as AIs grow more capable and appear on track to wield enormous cultural influence, directly advise on major government decisions, and even operate military equipment autonomously. We simply can’t tell what models, if any, should be trusted with such authority. Neel Nanda of Google DeepMind is one of the founding figures of the field of machine learning trying to fix this situation — mechanistic interpretabi

Neel Nanda on the race to read AI minds (part 1) | 80,000 Hours Search for: On this page: 1 Introduction 1.1 The episode in a nutshell 2 Highlights 3 Articles, books, and other media discussed in the show 4 Transcript 4.1 Cold open [00:00:00] 4.2 Who's Neel Nanda? [00:01:04] 4.3 How would mechanistic interpretability help with AGI [00:02:01] 4.4 What's mech interp? [00:05:12] 4.5 How Neel changed his take on mech interp [00:09:50] 4.6 Top successes in interpretability [00:16:00] 4.7 Probes can cheaply detect harmful intentions in AIs [00:20:13] 4.8 In some ways we understand AIs better than hu

Explore this link on the map →

related reading