✳flâneur — a map of the web's best reading
AI Chess Leaderboard - dubesor AI project
dubesor.de · 895 words · saved by 1 readers
LLM AI Chess Benchmark Leaderboard: Ranking, Elo, and Chess Performance of AI language models.
AI Chess Leaderboard - dubesor AI project ⚖️ 🌟 💭 ➡️ Mixed Why Chess? I like it. Plus, it's a historic centuries-old game of intellect, pure strategy with objective ground truth. Due to its exponential complexity beyond opening moves, it's largely resistant to common 'benchmaxxing' strategies. Tests game knowledge, reasoning, planning, state tracking, consistency and instruction adherence — measurable via objective superhuman judge (Stockfish) and updated with self-correcting Elo. It serves as a fantastic proxy, with rich metrics (Elo, accuracy, token efficiency, output speed, illegal outputs
Explore this link on the map →saved by
related reading
- LLM Chess: Benchmarking Reasoning and Instruction-Following in LLMs through Chessarxiv.org
- Puzzles | Paradigmparadigm.xyz
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Elo rating system - Wikipediaen.wikipedia.org
- Chess-GPT’s Internal World Model | Adam Karvonenadamkarvonen.github.io
- There's An AI For That® — The front page of AItheresanaiforthat.com
- AI Benchmark Leaderboards & Model Evals | BenchmarkListbenchmarklist.com
- Playing chess with large language modelsnicholas.carlini.com
- ML vs. Score matchingarxiv.org
- GitHub - open-compass/opencompass: OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets. · GitHubgithub.com
- Contra Labs - Powered by Contracontralabs.com
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com