flâneur — a map of the web's best reading

Sparsify: A mechanistic interpretability research agenda — AI Alignment Forum

alignmentforum.org · 11,393 words · saved by 1 readers

Over the last couple of years, mechanistic interpretability has seen substantial progress. Part of this progress has been enabled by the identification of superposition as a key barrier to understanding neural networks (Elhage et al., 2022) and the identification of sparse autoencoders as a solution to superposition (Sharkey et al., 2022; Cunningham et al., 2023; Bricken et al., 2023). From our current vantage point, I think there’s a relatively clear roadmap toward a world where mechanistic interpretability is useful for safety. This post outlines my views on what progress in mechanistic interpretability looks like and what I think is achievable by the field in the next 2+ years. It represents a rough outline of what I plan to work on in the near future. My thinking and work is, of course, very heavily inspired by the work of Chris Olah, other Anthropic researchers, and other early mechanistic interpretability researchers. In addition to sharing some personal takes, this article bring

x Sparsify: A mechanistic interpretability research agenda — AI Alignment Forum Interpretability (ML & AI) Sparse Autoencoders (SAEs) Research Agendas Apollo Research (org) Practice & Philosophy of Science AI Frontpage 43 Sparsify: A mechanistic interpretability research agenda by Lee Sharkey 3rd Apr 2024 27 min read 23 43 Over the last couple of years, mechanistic interpretability has seen substantial progress. Part of this progress has been enabled by the identification of superposition as a key barrier to understanding neural networks ( Elhage et al., 2022 ) and the identification of sparse

Explore this link on the map →

related reading