Perfectly Normal
Recently I was chatting with a friend, and described what I do in my current research using an analogy. This analogy seemed to be pretty helpful at clearing up some misconceptions about this research agenda and bridging inferential gaps, so I'm writing it up here. Note - this definitely glosses over some important aspects, but hopefully it communicates some of the key ideas. I usually start by describing my job as "neuroscience, for AI". Since AI (and in particular language models like ChatGPT) are getting increasingly powerful, it's becoming more important than ever to understand how they think. Ideally, we can improve our understanding of "AI brains" to a level that allows us detect things like deception or power-seeking behaviour. Essentially, we're hoping to build a "mind-reader for AI". Imagine if you were trying to build a mind-reader for humans, which worked by scanning human brains (and assume this is all the information you had access to, i.e. no reading expressions or taking
Perfectly Normal --> --> PERFECTLY NORMAL CALLUM MCDOUGALL This image was created using a variant of my thread art algorithm - read more here . What I Do For A Living (More Or Less) Recently I was chatting with a friend, and described what I do in my current research using an analogy. This analogy seemed to be pretty helpful at clearing up some misconceptions about this research agenda and bridging inferential gaps, so I'm writing it up here. Note - this definitely glosses over some important aspects, but hopefully it communicates some of the key ideas. Analogy I usually start by describing my
Explore this link on the map →saved by
related reading
- AI 2027ai-2027.com
- AI 2027ai-2027.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Next-Token Predictor Is An AI's Job, Not Its Speciesastralcodexten.com
- Language Models in Plato's Cave - by Sergey Levinesergeylevine.substack.com
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- What If We Had Bigger Brains? Imagining Minds beyond Ours-Stephen Wolfram Writingswritings.stephenwolfram.com
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- I Trained a Language Model. Then I Built a Brain Scanner and Looked Inside It. | by Caleb DeLeeuw | Mediummedium.com
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- On Optimism for Interpretabilitygoodfire.ai