flâneur — a map of the web's best reading

Induction heads - illustrated — LessWrong

lesswrong.com · 2,121 words · saved by 5 readers

This is my illustrated walkthrough of induction heads. I created it in order to concisely capture all the information about how the circuit works. There are 2 versions of the walkthrough: The final image from version 1 is inline below, and depending on your level of familiarity with transformers, looking at this diagram might provide most of the value of this post. If it doesn't make sense to you, then read on for the full walkthrough, where I build up this diagram bit by bit. Induction heads are a well-studied and understood circuit in transformers. They allow a model to perform in-context learning, of a very specific form: if a sequence contains a repeated subsequence e.g. of the form A B ... A B (where A and B stand for generic tokens, e.g. the first and last name of a person who doesn't appear in any of the model's training data), then the second time this subsequence occurs the transformer will be able to predict that B follows A. Although this might seem like weirdly specific abi

x Induction heads - illustrated — LessWrong Distillation & Pedagogy Has Diagram Interpretability (ML & AI) AI Frontpage 137 Induction heads - illustrated by CallumMcDougall 2nd Jan 2023 3 min read 14 137 Many thanks to everyone who provided helpful feedback, particularly Aryan Bhatt and Lawrence Chan! TL;DR This is my illustrated walkthrough of induction heads. I created it in order to concisely capture all the information about how the circuit works. There are 2 versions of the walkthrough: Version 1 is the one included in this post. It's slightly shorter, and focuses more on the intuitions t

Explore this link on the map →

saved by

related reading