Leela Interpretability
One of our early experiments was to do activation patching. We patch a small part of Leela's activations from the forward pass of a corrupted version of a puzzle into the forward pass on the original puzzle board state. Measuring the effect on the final output tells us how important that part of Leela's activations was.
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network 1UC Berkeley, 2Independent NeurIPS 2024 Choose A Chessboard State Do neural networks learn to implement algorithms involving look-ahead or search in the wild? Or do they only ever learn simple heuristics? We investigate this question for Leela Chess Zero, arguably the strongest existing chess-playing network. We find intriguing evidence of learned look-ahead in a single forward pass. This website showcases some of our results, see our paper for much more. Setup We consider chess puzzles such as the following (you can…
related reading
- Grandmaster-Level Chess Without Searcharxiv.org
- Chess-GPT’s Internal World Model | Adam Karvonenadamkarvonen.github.io
- OthelloGPT learned a bag of heuristics — LessWronglesswrong.com
- deepchess.pdfcs.tau.ac.il
- Transformer Progress | Leela Chess Zerolczero.org
- Manipulating Chess-GPT’s World Model | Adam Karvonenadamkarvonen.github.io
- As Rocks May Think | Eric Jangevjang.com
- Actually, Othello-GPT Has A Linear Emergent World Representation - Neel Nandaneelnanda.io
- Playing chess with large language modelsnicholas.carlini.com
- Learning to Play Chess from Textbooks (LEAP): a Corpus for Evaluating Chess Moves based on Sentiment Analysisarxiv.org
- Something weird is happening with LLMs and chessdynomight.substack.com
- A Very Unlikely Chess Game | Slate Star Codexslatestarcodex.com