Something weird is happening with LLMs and chess
A year ago, there was a lot of talk about large language models (LLMs) playing chess. Word was that if you trained a big enough model on enough text, then you could send it a partially played game, ask it to predict the next move, and it would play at the level of an advanced amateur. This seemed important. These are “language” models, after all, designed to predict language. Now, modern LLMs are trained on a sizeable fraction of all the text ever created. This surely includes many chess games. But they weren’t designed to be good at chess. And the games that are available are just lists of moves. Yet people found that LLMs could play all the way through to the end game, with never-before-seen boards. Did the language models build up some kind of internal representation of board state? And how to construct that state from lists of moves in chess’s extremely confusing notation? And how valuable different pieces and positions are? And how to force checkmate in an end-game? And they did t
A year ago, there was a lot of talk about large language models (LLMs) playing chess. Word was that if you trained a big enough model on enough text, then you could send it a partially played game, ask it to predict the next move, and it would play at the level of an advanced amateur. This seemed important. These are “language” models, after all, designed to predict language. Now, modern LLMs are trained on a sizeable fraction of all the text ever created. This surely includes many chess games. But they weren’t designed to be good at chess. And the games that are available are just lists of mo
Explore this link on the map →