flâneur — a map of the web's best reading

Chapter 1: Transformer Interpretability - ARENA

learn.arena.education · 1,864 words · saved by 1 readers

Please send any problems / bugs on the #errata channel in the Slack group, and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals, (1) Transformer Interpretability, (2) RL. This exercise set is built around linear probing, one of the most important tools in mechanistic interpretability for understanding what information language models represent internally. We'll look at three papers: The core idea: extract internal activations from a model, then train a simple classifier on them. If a linear probe can accurately classify some property from the activations, that property is linearly represented in the model's internal state. From the Geometry of Truth paper: "We identify a linear representation of truth that generalizes across several structurally and topically diverse datas

[1.3.1] Linear Probes Colab: exercises | solutions Please send any problems / bugs on the #errata channel in the Slack group , and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals , (1) Transformer Interpretability , (2) RL . Introduction This exercise set is built around linear probing , one of the most important tools in mechanistic interpretability for understanding what inform

Explore this link on the map →

saved by

related reading