What Would Non-Linear Features Actually Look Like? | Liv Gorton
“Non-linear representations” have become a catch-all objection to mechanistic interpretability work. The concern is worth taking seriously, but as typically stated, it collapses together cases with completely different implications and likelihoods.
What Would Non-Linear Features Actually Look Like? January 27, 2026 / 15 min read “Non-linear representations” have become a catch-all objection to mechanistic interpretability work. The concern is worth taking seriously, but as typically stated, it collapses together cases with completely different implications and likelihoods. Debates about non-linear features have often been unconstructive for several reasons. One common issue is that people have used the word “linear feature” to mean multiple different things, leading to miscommunication. Given this, the first contribution of this post is
Explore this link on the map →saved by
related reading
- Toy Models of Superpositiontransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Transformer Circuits Threadtransformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Interpretability Dreamstransformer-circuits.pub
- Circuits Updates — May 2023transformer-circuits.pub
- how neural networks think at scalemarmik.xyz
- Circuits Updates - January 2024transformer-circuits.pub
- Circuits Updates - July 2024transformer-circuits.pub
- Feature Visualizationdistill.pub