flâneur — a map of the web's best reading

The Rate Distortion Dance of Sparse Autoencoders | Tilde

tilderesearch.com · saved by 1 readers

Overview: in this blog post, we are going to be setting some of the theoretical foundations and intuition for the problems we think about. Over the coming week, we will release different blog posts focused on specific experiments and empirical questions. As such, this post aims to lay the groundwork for what's to come. We're excited to share the tip of the iceberg! The field of bottom-up mechanistic interpretability aims to elucidate the basic building blocks of machine cognition, which can be assembled hierarchically to build complex and emergent structures. In the early days of interpretability, researchers hypothesized that the fundamental unit of computation was the neuron and attempted to causally identify the role of individual neurons in early convolutional neural networks. These methods were met with limited success; researchers could occasionally cherry-pick one neuron or a set of neurons uniquely responsible for a task. However, the vast majority of the network did not appear

Overview: in this blog post, we are going to be setting some of the theoretical foundations and intuition for the problems we think about. Over the coming week, we will release different blog posts focused on specific experiments and empirical questions. As such, this post aims to lay the groundwork for what's to come. We're excited to share the tip of the iceberg! The field of bottom-up mechanistic interpretability aims to elucidate the basic building blocks of machine cognition, which can be assembled hierarchically to build complex and emergent structures. In the early days of interpretabil

Explore this link on the map →