flâneur — a map of the web's best reading

Gears-Level Mental Models of Transformer Interpretability — LessWrong

lesswrong.com · 2,207 words · saved by 1 readers

This post aims to quickly break down and explain the dominant mental models interpretability researchers currently use when thinking about how transf…

x Gears-Level Mental Models of Transformer Interpretability — LessWrong Interpretability (ML & AI) Language Models (LLMs) AI Frontpage 77 Gears-Level Mental Models of Transformer Interpretability by RowanWang 29th Mar 2022 AI Alignment Forum 7 min read 4 77 Ω 25 This post aims to quickly break down and explain the dominant mental models interpretability researchers currently use when thinking about how transformers work. In my view, the focus of transformer interpretability research is teleological: we care about the functions each component in the transformer performs and how those functions

Explore this link on the map →

related reading