✳flâneur — a map of the web's best reading
interpreting GPT: the logit lens — LessWrong
lesswrong.com · 7,456 words · saved by 1 readers
This post relates an observation I've made in my work with GPT-2, which I have not seen made elsewhere. …
x interpreting GPT: the logit lens — LessWrong GPT Machine Learning (ML) Gears-Level Interpretability (ML & AI) AI Frontpage 278 interpreting GPT: the logit lens by nostalgebraist 31st Aug 2020 AI Alignment Forum 13 min read 38 278 Ω 80 This post relates an observation I've made in my work with GPT-2, which I have not seen made elsewhere. IMO, this observation sheds a good deal of light on how the GPT-2/3/etc models (hereafter just "GPT") work internally. There is an accompanying Colab notebook which will let you interactively explore the phenomenon I describe here. [Edit: updated with another
Explore this link on the map →related reading
- interpreting GPT: the logit lens — AI Alignment Forumalignmentforum.org
- Google Colabcolab.research.google.com
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- GPT in 60 Lines of NumPy | Jay Modyjaykmody.com
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- What Is ChatGPT Doing … and Why Does It Work?-Stephen Wolfram Writingswritings.stephenwolfram.com
- microgptkarpathy.github.io
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Transformer Circuits Threadtransformer-circuits.pub
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education