flâneur — a map of the web's best reading

Do sparse autoencoders find "true features"? — LessWrong

lesswrong.com · 6,655 words · saved by 1 readers

Thanks to Joseph Bloom and James Oldfield for giving feedback on drafts which helped improve the post In this post I'll discuss an apparent limitation of sparse autoencoders (SAEs) in their current formulation as they are applied to discovering the latent features within AI models such as transformer-based LLMs. In brief, I'll cover the following: We intend for SAEs to discover the "true features" (a term I'm borrowing from Anthropic's SAE paper) used by the target model e.g. a transformer-based LLM. There isn't a universally accepted definition of what "true features" are, but for now I'll use the term somewhat loosely to refer to something like: There may be other ways of thinking about features but this should give us enough to work with for our current purposes. Consider a toy setup where one of the hidden layers in the target model has 3 "true features" represented by the following directions in its activation space: Additionally, suppose that feature 1 and feature 2 occur far mor

x Do sparse autoencoders find "true features"? — LessWrong Interpretability (ML & AI) Sparse Autoencoders (SAEs) AI Frontpage 76 Do sparse autoencoders find "true features"? by Demian Till 22nd Feb 2024 13 min read 33 76 Thanks to Joseph Bloom and James Oldfield for giving feedback on drafts which helped improve the post In this post I'll discuss an apparent limitation of sparse autoencoders (SAEs) in their current formulation as they are applied to discovering the latent features within AI models such as transformer-based LLMs. In brief, I'll cover the following: I'll argue that the L1 regula

Explore this link on the map →

related reading