Whole-History Rating: A Bayesian Rating System for Players of Time-Varying Strength
Whole-History Rating (WHR) is a new method to estimate the time-varying strengths of players involved in paired comparisons. Like many variations of the Elo rating system, the whole-history approach is based on the dynamic Bradley-Terry model. But, instead of using incremental approximations, WHR directly computes the exact maximum a posteriori over the whole rating history of all players. This additional accuracy comes at a higher computational cost than traditional methods, but computation is still fast enough to be easily applied in real time to large-scale game servers (a new game is added in less than 0.001 second). Experiments demonstrate that, in comparison to Elo, Glicko, TrueSkill, and decayed-history algorithms, WHR produces better predictions.
Whole-History Rating: A Bayesian Rating System for Players of Time-Varying Strength Whole-History Rating: A Bayesian Rating System for Players of Time-Varying Strength by Rémi Coulom Abstract Whole-History Rating (WHR) is a new method to estimate the time-varying strengths of players involved in paired comparisons. Like many variations of the Elo rating system, the whole-history approach is based on the dynamic Bradley-Terry model. But, instead of using incremental approximations, WHR directly computes the exact maximum a posteriori over the whole rating history of all players. This additional
Explore this link on the map →saved by
related reading
- Elo rating system - Wikipediaen.wikipedia.org
- Resorting Media Ratings · Gwern.netgwern.net
- How Not To Sort By Average Rating – Evan Millerevanmiller.org
- Bradley–Terry model - Wikipediaen.wikipedia.org
- [2207.00076] Efficient computation of rankings from pairwise comparisonsar5iv.labs.arxiv.org
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Shortform — LessWronglesswrong.com
- A statistical approach to model evaluations \ Anthropicanthropic.com
- Some data from LeelaPieceOdds — LessWronglesswrong.com
- Papers · Nikhil Garggargnikhil.com
- ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering | alphaXivalphaxiv.org
- RoboArena: Distributed Real-World Evaluation of Generalist Robot Policiesarxiv.org