AI Chess Leaderboard - dubesor AI project
dubesor.de · 895 words · saved by 1 readers
LLM AI Chess Benchmark Leaderboard: Ranking, Elo, and Chess Performance of AI language models.
AI Chess Leaderboard - dubesor AI project ⚖️ 🌟 💭 ➡️ Mixed Why Chess? I like it. Plus, it's a historic centuries-old game of intellect, pure strategy with objective ground truth. Due to its exponential complexity beyond opening moves, it's largely resistant to common 'benchmaxxing' strategies. Tests game knowledge, reasoning, planning, state tracking, consistency and instruction adherence — measurable via objective superhuman judge (Stockfish) and updated with self-correcting Elo. It serves as a fantastic proxy, with rich metrics (Elo, accuracy, token efficiency, output speed, illegal outputs
saved by
related reading
- LLM Chess: Benchmarking Reasoning and Instruction-Following in LLMs through Chessarxiv.org
- Puzzles | Paradigmparadigm.xyz
- BalatroBenchbalatrobench.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Chess-GPT’s Internal World Model | Adam Karvonenadamkarvonen.github.io
- Elo rating system - Wikipediaen.wikipedia.org
- Echo Chunk - Strategy games that generate themselvesechochunk.com
- BalatroBenchbalatrobench.com
- Grandmaster-Level Chess Without Searcharxiv.org
- Datacurve | The data engine for frontier AIdatacurve.ai
- FrontierSWEfrontierswe.com
- There's An AI For That® — The front page of AItheresanaiforthat.com