flâneur — a map of the web's best reading

Language models can explain neurons in language models

openaipublic.blob.core.windows.net · 758 words · saved by 5 readers

Language models have become more capable and more widely deployed, but we do not understand how they work. Recent work has made progress on understanding a small number of circuits and narrow behaviors, but to fully understand a language model, we'll need to analyze millions of neurons. This paper applies automation to the problem of scaling an interpretability technique to all the neurons in a large language model. Our hope is that building on this approach of automating interpretability will enable us to comprehensively audit the safety of models before deployment. Our technique seeks to explain what patterns in text cause a neuron to activate. It consists of three steps: This technique lets us leverage GPT-4 to define and automatically measure a quantitative notion of interpretability which we call an “explanation score”: a measure of a language model's ability to compress and reconstruct neuron activations using natural language.. The fact that this framework is quantitative allow

Language models can explain neurons in language models Language models can explain neurons in language models Authors Steven Bills ∗ , Nick Cammarata ∗ , Dan Mossing ∗ , Henk Tillman ∗ , Leo Gao ∗ , Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu ∗ , William Saunders ∗ * Core Research Contributor; Author contributions statement below . Correspondence to interpretability@openai.com . Affiliation OpenAI Published May 9, 2023 Contributions Methodology: Nick effectively started the project by having the initial idea to have GPT-4 explain neurons, and showing a simple explanation methodology worked

Explore this link on the map →

saved by

related reading