flâneur

Leela Interpretability

leela-interp.github.io · 894 words · saved by 1 readers

One of our early experiments was to do activation patching. We patch a small part of Leela's activations from the forward pass of a corrupted version of a puzzle into the forward pass on the original puzzle board state. Measuring the effect on the final output tells us how important that part of Leela's activations was.

Evidence of Learned Look-Ahead in a Chess-Playing Neural Network 1UC Berkeley, 2Independent NeurIPS 2024 Choose A Chessboard State Do neural networks learn to implement algorithms involving look-ahead or search in the wild? Or do they only ever learn simple heuristics? We investigate this question for Leela Chess Zero, arguably the strongest existing chess-playing network. We find intriguing evidence of learned look-ahead in a single forward pass. This website showcases some of our results, see our paper for much more. Setup We consider chess puzzles such as the following (you can…

related reading