Learned feature representations are biased by complexity, learning order, position, and more
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. lampinen@google.com \reportnumber Representation learning, and interpreting learned representations, are key areas of focus in machine learning and neuroscience. Both fields generally use representations as a means to understand or improve a system’s computations. In this work, however, we explore surprising dissociations between representation and computation that may pose challenges for such efforts. We create datasets in which we attempt to match the computational role that different features play, while manipulating other properties of the features or the data. We train various deep learning architectures to compute these multiple abstract features about their inputs. We find that their learned feature representations are systematically biased toward
\correspondingauthor lampinen@google.com \reportnumber Learned feature representations are biased by complexity, learning order, position, and more Andrew Kyle Lampinen Google DeepMind Stephanie C. Y. Chan Google DeepMind Katherine Hermann Google DeepMind Abstract Representation learning, and interpreting learned representations, are key areas of focus in machine learning and neuroscience. Both fields generally use representations as a means to understand or improve a system’s computations. In this work, however, we explore surprising dissociations between representation and computation that m
Explore this link on the map →related reading
- Toy Models of Superpositiontransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- What Would Non-Linear Features Actually Look Like? — Liv Gortonlivgorton.com
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Feature Visualizationdistill.pub
- Lucius Bushnaq's Shortform — LessWronglesswrong.com
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- What’s up with LLMs representing XORs of arbitrary features? — LessWronglesswrong.com
- Interpretability Dreamstransformer-circuits.pub
- Circuits Updates - January 2024transformer-circuits.pub
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com