About — Neel Nanda
Hi, I’m Neel! I’m interested in a lot of things, but especially maths, truth, improving myself and improving the world. I consider myself part of the Effective Altruism and rationality communities.
About Me Hi, I’m Neel! I run the Google DeepMind mechanistic interpretability team, our job is to take a trained neural network and try to reverse engineer the algorithms and structures it has learned. If you want to learn more about the field, see my appearance on the Machine Learning Street Talk podcast . I see the main goal of my work as reducing existential risk from AI, and I consider myself part of the Effective Altruism and rationality communities. Prior to this, I did independent mechanistic interpretability research, and I worked at Anthropic as a language model interpretability resea
Explore this link on the map →related reading
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- LessWronglesswrong.com
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Neel Nanda on the race to read AI minds (part 1) | 80,000 Hours80000hours.org
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- On Optimism for Interpretabilitygoodfire.ai
- Neel Nanda on leading a Google DeepMind team at 26 – and advice if you want to work at an AI company (part 2) | 80,000 Hours80000hours.org
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com
- What is the purpose of interpretability?ericjmichaud.com