flâneur — a map of the web's best reading

Towards Monosemanticity: Decomposing Language Models With Dictionary Learning

transformer-circuits.pub · 6,541 words · saved by 1 readers

Using a sparse autoencoder, we extract a large number of interpretable features from a one-layer transformer. Browse A/1 Features → Browse All Features → Not published yet. No DOI yet. Mechanistic interpretability seeks to understand neural networks by breaking them into components that are more easily understood than the whole. By understanding the function of each component, and how they interact, we hope to be able to reason about the behavior of the entire network. The first step in that program is to identify the correct components to analyze. Unfortunately, the most natural computational unit of the neural network – the neuron itself – turns out not to be a natural unit for human understanding. This is because many neurons are polysemantic: they respond to mixtures of seemingly unrelated inputs. In the vision model Inception v1, a single neuron responds to faces of cats and fronts of cars . In a small language model we discuss in this paper, a single neuron responds to a mixture

Towards Monosemanticity: Decomposing Language Models With Dictionary Learning Transformer Circuits Thread Towards Monosemanticity: Decomposing Language Models With Dictionary Learning Towards Monosemanticity: Decomposing Language Models With Dictionary Learning Using a sparse autoencoder, we extract a large number of interpretable features from a one-layer transformer. Browse A/1 Features → Browse All Features → Authors Trenton Bricken * , Adly Templeton * , Joshua Batson * , Brian Chen * , Adam Jermyn * , Tom Conerly, Nicholas L Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby,

Explore this link on the map →

related reading