✳flâneur — a map of the web's best reading
OthelloGPT learned a bag of heuristics — LessWrong
lesswrong.com · 5,177 words · saved by 1 readers
Work performed as a part of Neel Nanda's MATS 6.0 (Summer 2024) training program. …
x OthelloGPT learned a bag of heuristics — LessWrong MATS Program Interpretability (ML & AI) AI Frontpage 111 OthelloGPT learned a bag of heuristics by jylin04 , JackS , Adam Karvonen , Can 2nd Jul 2024 AI Alignment Forum 11 min read 10 111 Ω 43 Work performed as a part of Neel Nanda's MATS 6.0 (Summer 2024) training program. TLDR This is an interim report on reverse-engineering Othello-GPT , an 8-layer transformer trained to take sequences of Othello moves and predict legal moves. We find evidence that Othello-GPT learns to compute the board state using many independent decision rules that ar
Explore this link on the map →saved by
related reading
- Actually, Othello-GPT Has A Linear Emergent World Representation - Neel Nandaneelnanda.io
- Actually, Othello-GPT Has A Linear Emergent World Representation — AI Alignment Forumalignmentforum.org
- [2210.13382] Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Taskarxiv.org
- Othello-GPT: Reflections on the Research Process — LessWronglesswrong.com
- Othello-GPT: Reflections on the Research Process — LessWronglesswrong.com
- Chess-GPT’s Internal World Model | Adam Karvonenadamkarvonen.github.io
- Large Language Model: world models or surface statistics?thegradient.pub
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- A Very Unlikely Chess Game | Slate Star Codexslatestarcodex.com
- Playing chess with large language modelsnicholas.carlini.com
- Manipulating Chess-GPT’s World Model | Adam Karvonenadamkarvonen.github.io
- What Is ChatGPT Doing … and Why Does It Work?-Stephen Wolfram Writingswritings.stephenwolfram.com