flâneur — a map of the web's best reading

Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations — LessWrong

lesswrong.com · 7,303 words · saved by 1 readers

We present an information-theoretic framework for interpreting SAEs as lossy compression algorithms for communicating explanations of neural activations. We appeal to the Minimal Description Length (MDL) principle to motivate explanations of activations which are both accurate and concise.

x Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs — LessWrong MATS Program Sparse Autoencoders (SAEs) Causality Information theory Interpretability (ML & AI) Frontpage 43 Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs by Kola Ayonrinde , Michael Pearce , Lee Sharkey 23rd Aug 2024 19 min read 8 43 This work was produced as part of the ML Alignment & Theory Scholars Program - Summer 24 Cohort, under mentorship from Lee Sharkey and Jan Kulveit. Note: An updated paper version of this post can b

Explore this link on the map →

related reading