flâneur — a map of the web's best reading

Efficient Dictionary Learning with Switch Sparse Autoencoders — LessWrong

lesswrong.com · 5,799 words · saved by 1 readers

To recover all the relevant features from a superintelligent language model, we will likely need to scale sparse autoencoders (SAEs) to billions of features. Using current architectures, training extremely wide SAEs across multiple layers and sublayers at various sparsity levels is computationally intractable. Conditional computation has been used to scale transformers (Fedus et al.) to trillions of parameters while retaining computational efficiency. We introduce the Switch SAE, a novel architecture that leverages conditional computation to efficiently scale SAEs to many more features. The internal computations of large language models are inscrutable to humans. We can observe the inputs and the outputs, as well as every intermediate step in between, and yet, we have little to no sense of what the model is actually doing. For example, is the model inserting security vulnerabilities or backdoors into the code that it writes? Is the model lying, deceiving or seeking power? Deploying a s

x Efficient Dictionary Learning with Switch Sparse Autoencoders — LessWrong Interpretability (ML & AI) Sparse Autoencoders (SAEs) MATS Program Language Models (LLMs) AI Frontpage 118 Efficient Dictionary Learning with Switch Sparse Autoencoders by Anish Mudide 22nd Jul 2024 14 min read 20 118 Produced as part of the ML Alignment & Theory Scholars Program - Summer 2024 Cohort 0. Summary To recover all the relevant features from a superintelligent language model, we will likely need to scale sparse autoencoders (SAEs) to billions of features. Using current architectures, training extremely wide

Explore this link on the map →

related reading