Towards Automated Circuit Discovery for Mechanistic Interpretability
arxiv.org · 7,455 words · saved by 2 readers
N/A
Towards Automated Circuit Discovery for Mechanistic Interpretability Arthur Conmy∗ Augustine N. Mavor-Parker∗ Aengus Lynch∗ Stefan Heimersheim Independent UCL UCL University of Cambridge arXiv:2304.14997v4 [cs.LG] 28 Oct 2023 Adrià Garriga-Alonso∗…
saved by
related reading
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] — LessWronglesswrong.com
- Transformer Circuits Threadtransformer-circuits.pub
- Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decompositionarxiv.org
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Mechanistic Interpretability: Circuits, Induction Headsmbrenndoerfer.com
- Softmax Linear Unitstransformer-circuits.pub
- The Building Blocks of Interpretabilitydistill.pub
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io