Chapter 1: Transformer Interpretability - ARENA
Please send any problems / bugs on the #errata channel in the Slack group, and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals, (1) Transformer Interpretability, (2) RL. Note - unless otherwise specified, first person here refers to the primary researcher, Neel Nanda. Emergent World Representations is a fascinating recent ICLR Oral paper from Kenneth Li et al, summarised in Kenneth's excellent post on the Gradient. They trained a model (Othello-GPT) to play legal moves in the board game Othello, by giving it random games (generated by choosing a legal next move uniformly at random) and training it to predict the next move. The headline result is that Othello-GPT learns an emergent world representation - despite never being explicitly given the state of the board, and just
[1.5.3] OthelloGPT Colab: exercises | solutions Please send any problems / bugs on the #errata channel in the Slack group , and ask any questions on the dedicated channels for this chapter of material. If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme. Links to all other chapters: (0) Fundamentals , (1) Transformer Interpretability , (2) RL . Introduction Note - unless otherwise specified, first person here refers to the primary researcher, Neel Nanda. Emergent World Representations is a fascinating
Explore this link on the map →related reading
- Actually, Othello-GPT Has A Linear Emergent World Representation - Neel Nandaneelnanda.io
- Actually, Othello-GPT Has A Linear Emergent World Representation — AI Alignment Forumalignmentforum.org
- [2210.13382] Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Taskarxiv.org
- Structure and Interpretation of Deep Networkssidn.baulab.info
- Othello-GPT: Reflections on the Research Process — LessWronglesswrong.com
- Othello-GPT: Reflections on the Research Process — LessWronglesswrong.com
- Transformer Circuits Threadtransformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- OthelloGPT learned a bag of heuristics — LessWronglesswrong.com
- Chess-GPT’s Internal World Model | Adam Karvonenadamkarvonen.github.io
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education