✳flâneur — a map of the web's best reading
Leela Interpretability
leela-interp.github.io · saved by 1 readers
One of our early experiments was to do activation patching. We patch a small part of Leela's activations from the forward pass of a corrupted version of a puzzle into the forward pass on the original puzzle board state. Measuring the effect on the final output tells us how important that part of Leela's activations was.
Explore this link on the map →related reading
- Attribution Patching: Activation Patching At Industrial Scale - Neel Nandaneelnanda.io
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- [2404.15255] How to use and interpret activation patchingarxiv.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Othello-GPT: Reflections on the Research Process — LessWronglesswrong.com
- Othello-GPT: Reflections on the Research Process — LessWronglesswrong.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Exploratory Analysis Demo - TransformerLens Documentationtransformerlensorg.github.io
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWronglesswrong.com