OthelloGPT learned a bag of heuristics — LessWrong
lesswrong.com · 5,177 words · saved by 1 readers
Work performed as a part of Neel Nanda's MATS 6.0 (Summer 2024) training program. …
x OthelloGPT learned a bag of heuristics — LessWrong MATS Program Interpretability (ML & AI) AI Frontpage 111 OthelloGPT learned a bag of heuristics by jylin04 , JackS , Adam Karvonen , Can 2nd Jul 2024 AI Alignment Forum 11 min read 10 111 Ω 43 Work performed as a part of Neel Nanda's MATS 6.0 (Summer 2024) training program. TLDR This is an interim report on reverse-engineering Othello-GPT , an 8-layer transformer trained to take sequences of Othello moves and predict legal moves. We find evidence that Othello-GPT learns to compute the board state using many independent decision rules that ar
saved by
related reading
- Actually, Othello-GPT Has A Linear Emergent World Representation - Neel Nandaneelnanda.io
- interpreting GPT: the logit lens — LessWronglesswrong.com
- [2210.13382] Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Taskarxiv.org
- Actually, Othello-GPT Has A Linear Emergent World Representation — AI Alignment Forumalignmentforum.org
- Chess-GPT’s Internal World Model | Adam Karvonenadamkarvonen.github.io
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Taskarxiv.org
- Othello-GPT: Reflections on the Research Process — LessWronglesswrong.com
- Othello-GPT: Reflections on the Research Process — LessWronglesswrong.com
- Large Language Model: world models or surface statistics?thegradient.pub
- As Rocks May Think | Eric Jangevjang.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Structure and Interpretation of Deep Networkssidn.baulab.info