Structure and Interpretation of Deep Networks
The authors of the papers of today's discussion are mainly Kenneth Li, PhD student at Harvard University, and Dr. Yanai Elazar is postdoctoral researcher on the AllenNLP team at AI2. Kenneth Li is working on LLM dialogues and interpretability for alignment of LLMs. Dr. Yanai Elazar works on interpretability of generative model and has authored a great paper "Null it out: Guarding protected attributes by iterative nullspace projection" that showed how mathematically feature space could be made unresponsive by simple linear algebra. Probing is an attempt by computer scientists to understand the workings of neural networks. The most popular way of probing is by learning to make sense of a representation of a neural network by keeping the information in its purest form as much as possible. The information under scrutiny is usually a human interpretable property of known data that is believed could be decoded in the easiest way if the information is present in the representation and also in
Structure and Interpretation of Deep Networks Probing September 19, 2024 • Rahul Chowdhury, Ritik Bompilwar Who are the paper authors? The authors of the papers of today's discussion are mainly Kenneth Li, PhD student at Harvard University, and Dr. Yanai Elazar is postdoctoral researcher on the AllenNLP team at AI2. Kenneth Li is working on LLM dialogues and interpretability for alignment of LLMs. Dr. Yanai Elazar works on interpretability of generative model and has authored a great paper "Null it out: Guarding protected attributes by iterative nullspace projection" that showed how mathematic
Explore this link on the map →related reading
- Actually, Othello-GPT Has A Linear Emergent World Representation - Neel Nandaneelnanda.io
- [2210.13382] Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Taskarxiv.org
- Actually, Othello-GPT Has A Linear Emergent World Representation — AI Alignment Forumalignmentforum.org
- Large Language Model: world models or surface statistics?thegradient.pub
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org