flâneur — a map of the web's best reading

Addressing Feature Suppression in SAEs — LessWrong

lesswrong.com · 3,498 words · saved by 1 readers

Produced as part of the ML Alignment Theory Scholars Program - Winter 2023-24 Cohort as part of Lee Sharkey's stream. …

x Addressing Feature Suppression in SAEs — LessWrong Sparse Autoencoders (SAEs) Interpretability (ML & AI) AI Frontpage 88 Addressing Feature Suppression in SAEs by Benjamin Wright , Lee Sharkey 16th Feb 2024 AI Alignment Forum 12 min read 5 88 Ω 46 Produced as part of the ML Alignment Theory Scholars Program - Winter 2023-24 Cohort as part of Lee Sharkey's stream. TL;DR Sparse autoencoders are a method of resolving superposition by recovering linearly encoded “features” inside activations. Unfortunately, despite the great recent success of SAEs at extracting human interpretable features, they

Explore this link on the map →

related reading