Tiny Mech Interp Projects: Emergent Positional Embeddings of Words — Neel Nanda
This post was written in a rush and represents a few hours of research on a thing I was curious about, and is an exercise in being less of a perfectionist. I'd love to see someone build on this work! Thanks a lot to Wes Gurnee for pairing with me on this Tokens are weird, man A particularly notable observation in interpretability in the wild is that part of the studied circuit moves around information about whether the indirect object of the sentence is the first or second name in the sentence. The natural guess is that heads are moving around the absolute position of the correct name. But even in prompt formats where the first and second names are in the different absolute positions, they find that the informations conveyed by these heads are exactly the same, and can be patched between prompt templates! (credit to Alexandre Variengien for making this point to me). This raises the possibility that the model has learned what I call emergent positional embeddings - rather than represen
Tiny Mech Interp Projects: Emergent Positional Embeddings of Words Jul 18 Written By Neel Nanda This post was written in a rush and represents a few hours of research on a thing I was curious about, and is an exercise in being less of a perfectionist. I'd love to see someone build on this work! Thanks a lot to Wes Gurnee for pairing with me on this Tokens are weird, man Introduction A particularly notable observation in interpretability in the wild is that part of the studied circuit moves around information about whether the indirect object of the sentence is the first or second name in the s
Explore this link on the map →saved by
related reading
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Extending Context is Hard | kaiokendevkaiokendev.github.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- GPT-2's positional embedding matrix is a helix — LessWronglesswrong.com
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Fact Finding: Simplifying the Circuit (Post 2) — LessWronglesswrong.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education