Language models can explain neurons in language models
Language models have become more capable and more widely deployed, but we do not understand how they work. Recent work has made progress on understanding a small number of circuits and narrow behaviors, but to fully understand a language model, we'll need to analyze millions of neurons. This paper applies automation to the problem of scaling an interpretability technique to all the neurons in a large language model. Our hope is that building on this approach of automating interpretability will enable us to comprehensively audit the safety of models before deployment. Our technique seeks to explain what patterns in text cause a neuron to activate. It consists of three steps: This technique lets us leverage GPT-4 to define and automatically measure a quantitative notion of interpretability which we call an “explanation score”: a measure of a language model's ability to compress and reconstruct neuron activations using natural language.. The fact that this framework is quantitative allow
Language models can explain neurons in language models Language models can explain neurons in language models Authors Steven Bills ∗ , Nick Cammarata ∗ , Dan Mossing ∗ , Henk Tillman ∗ , Leo Gao ∗ , Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu ∗ , William Saunders ∗ * Core Research Contributor; Author contributions statement below . Correspondence to interpretability@openai.com . Affiliation OpenAI Published May 9, 2023 Contributions Methodology: Nick effectively started the project by having the initial idea to have GPT-4 explain neurons, and showing a simple explanation methodology worked
Explore this link on the map →saved by
related reading
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Neuronpedianeuronpedia.org
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Interfaces for Explaining Transformer Language Models – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Language Models, World Models, and Human Model-Buildinglingo.csail.mit.edu
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com