flâneur — a map of the web's best reading

AI Chess Leaderboard - dubesor AI project

dubesor.de · 895 words · saved by 1 readers

LLM AI Chess Benchmark Leaderboard: Ranking, Elo, and Chess Performance of AI language models.

AI Chess Leaderboard - dubesor AI project ⚖️ 🌟 💭 ➡️ Mixed Why Chess? I like it. Plus, it's a historic centuries-old game of intellect, pure strategy with objective ground truth. Due to its exponential complexity beyond opening moves, it's largely resistant to common 'benchmaxxing' strategies. Tests game knowledge, reasoning, planning, state tracking, consistency and instruction adherence — measurable via objective superhuman judge (Stockfish) and updated with self-correcting Elo. It serves as a fantastic proxy, with rich metrics (Elo, accuracy, token efficiency, output speed, illegal outputs

Explore this link on the map →

saved by

related reading