✳flâneur — a map of the web's best reading
interpreting GPT: the logit lens - AI Alignment Forum
alignmentforum.org · 5,358 words · saved by 1 readers
This post relates an observation I've made in my work with GPT-2, which I have not seen made elsewhere. …
x interpreting GPT: the logit lens — AI Alignment Forum GPT Machine Learning (ML) Gears-Level Interpretability (ML & AI) AI Frontpage 80 interpreting GPT: the logit lens by nostalgebraist 31st Aug 2020 13 min read 38 80 This post relates an observation I've made in my work with GPT-2, which I have not seen made elsewhere. IMO, this observation sheds a good deal of light on how the GPT-2/3/etc models (hereafter just "GPT") work internally. There is an accompanying Colab notebook which will let you interactively explore the phenomenon I describe here. [Edit: updated with another section on compa
Explore this link on the map →related reading
- interpreting GPT: the logit lens — LessWronglesswrong.com
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- GPT in 60 Lines of NumPy | Jay Modyjaykmody.com
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Google Colabcolab.research.google.com
- What Is ChatGPT Doing … and Why Does It Work?-Stephen Wolfram Writingswritings.stephenwolfram.com
- microgptkarpathy.github.io
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education