Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations — LessWrong
We present an information-theoretic framework for interpreting SAEs as lossy compression algorithms for communicating explanations of neural activations. We appeal to the Minimal Description Length (MDL) principle to motivate explanations of activations which are both accurate and concise.
x Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs — LessWrong MATS Program Sparse Autoencoders (SAEs) Causality Information theory Interpretability (ML & AI) Frontpage 43 Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs by Kola Ayonrinde , Michael Pearce , Lee Sharkey 23rd Aug 2024 19 min read 8 43 This work was produced as part of the ML Alignment & Theory Scholars Program - Summer 24 Cohort, under mentorship from Lee Sharkey and Jan Kulveit. Note: An updated paper version of this post can b
Explore this link on the map →related reading
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- pdfopenreview.net
- An Intuitive Explanation of Sparse Autoencoders for Mechanistic Interpretability of LLMs — LessWronglesswrong.com
- Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning — LessWronglesswrong.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- Activation space interpretability may be doomed — LessWronglesswrong.com
- Sparsify: A mechanistic interpretability research agenda — AI Alignment Forumalignmentforum.org
- A List of 45+ Mech Interp Project Ideas from Apollo Research’s Interpretability Team — AI Alignment Forumalignmentforum.org