Bridging the Attention Gap: Complete Replacement Models for Complete Circuit Tracing
We introduce Complete Replacement Models (CRMs), which combine transcoders for MLPs with Low-Rank Sparse Attention (Lorsa) modules that decompose attention into interpretable features.
← OpenMOSS Interpretability Posts We introduce Complete Replacement Models (CRMs), which combine transcoders for MLPs with Low-Rank Sparse Attention (Lorsa) modules that decompose attention into interpretable features. This enables us to build attribution graphs that reveal interpretable sparse computational paths in both attention and MLP computation. Architectural completeness also unlocks accurate and efficient global weight analysis, letting us construct global circuits that reveal input-independent versions of canonical circuits like induction at the feature level. OpenMOSS Team,…
saved by
related reading
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- On the Biology of a Large Language Modeltransformer-circuits.pub
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- How LLMs Actually Work | 0xkato0xkato.xyz
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Interpreting Language Model Parametersgoodfire.ai
- Sparse Attention Post-Training for Mechanistic Interpretabilityarxiv.org
- [2406.11944] Transcoders Find Interpretable LLM Feature Circuitsarxiv.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Language Model Circuits Are Sparse in the Neuron Basis | Transluce AItransluce.org