flâneur — a map of the web's best reading

[2409.14507] A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders

ar5iv.labs.arxiv.org · 8,945 words · saved by 1 readers

Sparse Autoencoders (SAEs) have emerged as a promising approach to decompose the activations of Large Language Models (LLMs) into human-interpretable latents. In this paper, we pose two questions. First, to what extent…

A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders David Chanin 1,2,* , James Wilken-Smith 1,* , Tomáš Dulka 1,* , Hardik Bhatnagar 1,* , Joseph Bloom 1,3 1 LASR Labs, 2 University College London, 3 Decode Research * These authors contributed equally to this work Abstract Sparse Autoencoders (SAEs) have emerged as a promising approach to decompose the activations of Large Language Models (LLMs) into human-interpretable latents. In this paper, we pose two questions. First, to what extent do SAEs extract monosemantic and interpretable latents? Second, to what e

Explore this link on the map →

related reading