flâneur — a map of the web's best reading

An Analogy for Understanding Transformers — LessWrong

lesswrong.com · 3,608 words · saved by 1 readers

Thanks to the following people for feedback: Tilman Rauker, Curt Tigges, Rudolf Laine, Logan Smith, Arthur Conmy, Joseph Bloom, Rusheb Shah, James Dao. I present an analogy for the transformer architecture: each vector in the residual stream is a person standing in a line, who is holding a token, and trying to guess what token the person in front of them is holding. Attention heads represent questions that people in this line can ask to everyone standing behind them (queries are the questions, keys determine who answers the questions, values determine what information gets passed back to the original question-asker), and MLPs represent the internal processing done by each person in the line. I claim this is a useful way to intuitively understand the transformer architecture, and I'll present several reasons for this (as well as ways induction heads and indirect object identification can be understood in these terms).[1] In this post, I'm going to present an analogy for understanding ho

x An Analogy for Understanding Transformers — LessWrong Transformer Circuits Transformers AI Frontpage 92 An Analogy for Understanding Transformers by CallumMcDougall 13th May 2023 11 min read 6 92 Thanks to the following people for feedback: Tilman Rauker, Curt Tigges, Rudolf Laine, Logan Smith, Arthur Conmy, Joseph Bloom, Rusheb Shah, James Dao. TL;DR I present an analogy for the transformer architecture: each vector in the residual stream is a person standing in a line, who is holding a token, and trying to guess what token the person in front of them is holding. Attention heads represent q

Explore this link on the map →

related reading