✳flâneur — a map of the web's best reading
Othello-GPT: Reflections on the Research Process — LessWrong
lesswrong.com · 4,471 words · saved by 1 readers
This is the third in a three post sequence about interpreting Othello-GPT. See the first post for context. …
x Othello-GPT: Reflections on the Research Process — LessWrong Interpreting Othello-GPT Interpretability (ML & AI) AI Frontpage 38 Othello-GPT: Reflections on the Research Process by Neel Nanda 29th Mar 2023 AI Alignment Forum Linkpost for neelnanda.io 18 min read 0 38 Ω 14 This is the third in a three post sequence about interpreting Othello-GPT. See the first post for context. This post is a detailed account of what my research process was, decisions made at each point, what intermediate results looked like, etc. It's deliberately moderately unpolished, in the hopes that it makes this more u
Explore this link on the map →related reading
- Othello-GPT: Reflections on the Research Process — LessWronglesswrong.com
- Actually, Othello-GPT Has A Linear Emergent World Representation - Neel Nandaneelnanda.io
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- OthelloGPT learned a bag of heuristics — LessWronglesswrong.com
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Actually, Othello-GPT Has A Linear Emergent World Representation — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Transformer Circuits Threadtransformer-circuits.pub
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- [2210.13382] Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Taskarxiv.org
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com