interpreting GPT: the logit lens — LessWrong
lesswrong.com · 7,456 words · saved by 9 readers
This post relates an observation I've made in my work with GPT-2, which I have not seen made elsewhere. …
x interpreting GPT: the logit lens — LessWrong GPT Machine Learning (ML) Gears-Level Interpretability (ML & AI) AI Frontpage 278 interpreting GPT: the logit lens by nostalgebraist 31st Aug 2020 AI Alignment Forum 13 min read 38 278 Ω 80 This post relates an observation I've made in my work with GPT-2, which I have not seen made elsewhere. IMO, this observation sheds a good deal of light on how the GPT-2/3/etc models (hereafter just "GPT") work internally. There is an accompanying Colab notebook which will let you interactively explore the phenomenon I describe here. [Edit: updated with another
saved by
- Shubham Shah
- Karan MJ
- Claire Wang
- Nicholas Charette
- Lydia Nottingham
- Timothy Kostolansky
- Cheikh Fiteni
- Julian H
- Chris Shi
related reading
- interpreting GPT: the logit lens — AI Alignment Forumalignmentforum.org
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Google Colabcolab.research.google.com
- GPT in 60 Lines of NumPy | Jay Modyjaykmody.com
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- What Is ChatGPT Doing … and Why Does It Work?-Stephen Wolfram Writingswritings.stephenwolfram.com
- Actually, Othello-GPT Has A Linear Emergent World Representation - Neel Nandaneelnanda.io
- microgptkarpathy.github.io
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education