flâneur — a map of the web's best reading

Sparse Autoencoders Reveal Universal Feature Spaces Across Large Language Models

arxiv.org · 18,483 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. We investigate feature universality in large language models (LLMs), a research field that aims to understand how different models similarly represent concepts in the latent spaces of their intermediate layers. Demonstrating feature universality allows discoveries about latent representations to generalize across several models. However, comparing features across LLMs is challenging due to polysemanticity, in which individual neurons often correspond to multiple features rather than distinct ones. This makes it difficult to disentangle and match features across different models. To address this issue, we employ a method known as dictionary learning by using sparse autoencoders (SAEs) to transform LLM activations into more interpretable spaces spanned by

Sparse Autoencoders Reveal Universal Feature Spaces Across Large Language Models Michael Lan Philip Torr ‡ Austin Meek § Ashkan Khakzar ‡ David Krueger ♡ Fazl Barez †‡ † Tangentic ‡ University of Oxford § University of Delaware ♡ MILA Work done during the ERA-Krueger AI Safety Lab internship. Author contributions detailed in § Author Contributions . Abstract We investigate feature universality in large language models (LLMs), a research field that aims to understand how different models similarly represent concepts in the latent spaces of their intermediate layers. Demonstrating feature univer

Explore this link on the map →

related reading