flâneur — a map of the web's best reading

Activation space interpretability may be doomed — LessWrong

lesswrong.com · 8,502 words · saved by 1 readers

TL;DR: There may be a fundamental problem with interpretability work that attempts to understand neural networks by decomposing their individual acti…

x Activation space interpretability may be doomed — LessWrong Apollo Research (org) Interpretability (ML & AI) AI Frontpage 2025 Top Fifty: 12 % 154 Activation space interpretability may be doomed by bilalchughtai , Lucius Bushnaq 8th Jan 2025 AI Alignment Forum 10 min read 34 154 Ω 58 TL;DR: There may be a fundamental problem with interpretability work that attempts to understand neural networks by decomposing their individual activation spaces in isolation: It seems likely to find features of the activations - features that help explain the statistical structure of activation spaces, rather

Explore this link on the map →

related reading