LLM Chess: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. We introduce LLM Chess, an evaluation framework designed to probe the generalization of reasoning and instruction-following abilities in large language models (LLMs) through extended agentic interaction in the domain of chess. We rank over 50 open and closed source models by playing against a random opponent using a range of behavioral metrics, including win and loss rates, m
LLM Chess : Benchmarking Reasoning and Instruction-Following in LLMs through Chess Sai Kolasani 1 , Maxim Saplin 2 , Nicholas Crispino 3 , Kyle Montgomery 3 , Jared Quincy Davis 4 , Matei Zaharia 1 , Chi Wang 5 , Chenguang Wang 3 1 UC Berkeley, 2 Independent Researcher, 3 UC Santa Cruz, 4 Stanford University, 5 Google DeepMind saikolasani@berkeley.edu , tutehabre@gmail.com Abstract We introduce LLM Chess , an evaluation framework designed to probe the generalization of reasoning and instruction-following abilities in large language models (LLMs) through extended agentic interaction in the doma
saved by
related reading
- AI Chess Leaderboard - dubesor AI projectdubesor.de
- Something weird is happening with LLMs and chessdynomight.substack.com
- Playing chess with large language modelsnicholas.carlini.com
- As Rocks May Think | Eric Jangevjang.com
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- Grandmaster-Level Chess Without Searcharxiv.org
- Chess-GPT’s Internal World Model | Adam Karvonenadamkarvonen.github.io
- DeepSeek-R1arxiv.org
- I Built TetrisBench, Where LLMs Compete at Playing Tetris. Here’s What I Found.a16z.com
- Manipulating Chess-GPT’s World Model | Adam Karvonenadamkarvonen.github.io
- Learning Chess With Language Models and Transformersarxiv.org
- Large Language Model: world models or surface statistics?thegradient.pub