Perfectly Normal
Recently I was chatting with a friend, and described what I do in my current research using an analogy. This analogy seemed to be pretty helpful at clearing up some misconceptions about this research agenda and bridging inferential gaps, so I'm writing it up here. Note - this definitely glosses over some important aspects, but hopefully it communicates some of the key ideas. I usually start by describing my job as "neuroscience, for AI". Since AI (and in particular language models like ChatGPT) are getting increasingly powerful, it's becoming more important than ever to understand how they think. Ideally, we can improve our understanding of "AI brains" to a level that allows us detect things like deception or power-seeking behaviour. Essentially, we're hoping to build a "mind-reader for AI". Imagine if you were trying to build a mind-reader for humans, which worked by scanning human brains (and assume this is all the information you had access to, i.e. no reading expressions or taking
Perfectly Normal --> --> PERFECTLY NORMAL CALLUM MCDOUGALL This image was created using a variant of my thread art algorithm - read more here . What I Do For A Living (More Or Less) Recently I was chatting with a friend, and described what I do in my current research using an analogy. This analogy seemed to be pretty helpful at clearing up some misconceptions about this research agenda and bridging inferential gaps, so I'm writing it up here. Note - this definitely glosses over some important aspects, but hopefully it communicates some of the key ideas. Analogy I usually start by describing my
saved by
related reading
- AI 2027ai-2027.com
- AI 2027ai-2027.com
- Mechanistic Interpretability: A Challenge Common to Both Artificial and Biological Intelligencekempnerinstitute.harvard.edu
- God Help Us, Let's Try To Understand The Paper On AI Monosemanticityastralcodexten.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Next-Token Predictor Is An AI's Job, Not Its Speciesastralcodexten.com
- How AI Is Learning to Think in Secretnickandresen.substack.com
- What is the purpose of interpretability?ericjmichaud.com
- Language Models in Plato's Cave - by Sergey Levinesergeylevine.substack.com
- What If We Had Bigger Brains? Imagining Minds beyond Ours-Stephen Wolfram Writingswritings.stephenwolfram.com
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org