Fetching the paper…
Reading the bibliography…
The Elo rating system is widely adopted to evaluate the skills of (chess) game and sports players.
OpenSpiel: A framework for reinforcement learning in games
Lanctot, M.; Lockhart, E.; Lespiau, J.-B.; Zambaldi, V.; Upadhyay, S.; Pérolat, J.; Srinivasan, S.; Timbers, F.; Tuyls, K.; Omidshafiei, S.; et al. 2019 · 1908
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R. 1933 · 1933
Earlier work this paper cites.
The rating of chessplayers, past and present
Elo, A. E. 1978 · 1978
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P.; Cesa-Bianchi, N.; and Fischer, P. 2002 · 2002
Earlier work this paper cites.
Regret Minimization in Stochastic Contextual Dueling Bandits
Saha, A.; and Gopalan, A. 2020 · 2002
Earlier work this paper cites.
MM algorithms for generalized Bradley-Terry models
Hunter, D. R.; et al. 2004 · 2004
Earlier work this paper cites.
Round robin scheduling–a survey
Rasmussen, R. V.; and Trick, M. A. 2008 · 2008
Earlier work this paper cites.
On the local optimality of lambdarank
Donmez, P.; Svore, K. M.; and Burges, C. J. 2009 · 2009
Earlier work this paper cites.
Interactively optimizing information retrieval systems as a dueling bandits problem
Yue, Y.; and Joachims, T. 2009 · 2009
Earlier work this paper cites.
Monte Carlo tree search in Hex
Arneson, B.; Hayward, R. B.; and Henderson, P. 2010 · 2010
Earlier work this paper cites.
Pure exploration in finitely-armed and continuous-armed bandits
Bubeck, S.; Munos, R.; and Stoltz, G. 2011 · 2011
Earlier work this paper cites.
Statistical ranking and combinatorial Hodge theory
Jiang, X.; Lim, L.-H.; Yao, Y.; and Ye, Y. 2011 · 2011
Earlier work this paper cites.
Analysis of thompson sampling for the multi-armed bandit problem
Agrawal, S.; and Goyal, N. 2012 · 2012
Earlier work this paper cites.
Optimal seedings in elimination tournaments
Groh, C.; Moldovanu, B.; Sela, A.; and Sunde, U. 2012 · 2012
Cited alongside, same era.
The k-armed dueling bandits problem
Yue, Y.; Broder, J.; Kleinberg, R.; and Joachims, T. 2012 · 2012
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Agrawal, S.; and Goyal, N. 2013 · 2013
Cited alongside, same era.
Trirank: Review-aware explainable recommendation by modeling aspects
He, X.; Chen, T.; Kan, M.-Y.; and Chen, X. 2015 · 2015
Cited alongside, same era.
Giraffe: Using deep reinforcement learning to play chess
Lai, M. 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; and Ostrovski, G. 2015 · 2015
Active ranking from pairwise comparisons and when parametric assumptions do not help
Heckel, R.; Shah, N. B.; Ramchandran, K.; Wainwright, M. J.; et al. 2019 · 2019
Later among the works it cites.
α \alpha -rank: Multi-agent evaluation by evolution
Omidshafiei, S.; Papadimitriou, C.; Piliouras, G.; Tuyls, K.; Rowland, M.; Lespiau, J.-B.; Czarnecki, W. M.; Lanctot, M.; Perolat, J.; and Munos, R. 2019 · 2019
Later among the works it cites.
Multiagent evaluation under incomplete information
Rowland, M.; Omidshafiei, S.; Tuyls, K.; Perolat, J.; Valko, M.; Piliouras, G.; and Munos, R. 2019 · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O.; Babuschkin, I.; Czarnecki, W. M.; Mathieu, M.; Dudzik, A.; Chung, J.; Choi, D. H.; Powell, R.; Ewalds, T.; Georgiev, P.; et al. 2019 · 2019
Later among the works it cites.
Real World Games Look Like Spinning Tops
Czarnecki, W. M.; Gidel, G.; Tracey, B.; Tuyls, K.; Omidshafiei, S.; Balduzzi, D.; and Jaderberg, M. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Online rank elicitation for plackett-luce: A dueling bandits approach
Szörényi, B.; Busa-Fekete, R.; Paul, A.; and Hüllermeier, E. 2015 · 2015
Cited alongside, same era.
Mastering the game of Go without human knowledge
Silver, D.; Schrittwieser, J.; Simonyan, K.; Antonoglou, I.; Huang, A.; Guez, A.; Hubert, T.; Baker, L.; Lai, M.; and Bolton, A. 2017 · 2017
Cited alongside, same era.
Re-evaluating evaluation
Balduzzi, D.; Tuyls, K.; Perolat, J.; and Graepel, T. 2018 · 2018
Cited alongside, same era.
The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning
Gruslys, A.; Dabney, W.; Azar, M. G.; Piot, B.; Bellemare, M.; and Munos, R. 2018 · 2018
Cited alongside, same era.
LIIR: learning individual intrinsic reward in multi-agent reinforcement learning
Du, Y.; Han, L.; Fang, M.; Dai, T.; Liu, J.; and Tao, D. 2019 · 2019
Cited alongside, same era.
Grid-wise control for multi-agent reinforcement learning in video game ai
Han, L.; Sun, P.; Du, Y.; Xiong, J.; Wang, Q.; Sun, X.; Liu, H.; and Zhang, T. 2019 · 2019
Cited alongside, same era.
A Generalized Training Approach for Multiagent Learning
Muller, P.; Omidshafiei, S.; Rowland, M.; Tuyls, K.; Perolat, J.; Liu, S.; Hennes, D.; Marris, L.; Lanctot, M.; Hughes, E.; et al. 2020 · 2020
Later among the works it cites.
α α \alpha^{\alpha} -Rank: Practically Scaling α \alpha -Rank through Stochastic Optimisation
Yang, Y.; Tutunov, R.; Sakulwongtana, P.; and Ammar, H. B. 2020 · 2020
Later among the works it cites.
An efficient algorithm for generalized linear bandit: Online stochastic gradient descent and thompson sampling
Ding, Q.; Hsieh, C.-J.; and Sharpnack, J. 2021 · 2021
Later among the works it cites.
Estimating α \alpha -Rank from A Few Entries with Low Rank Matrix Completion
Du, Y.; Yan, X.; Chen, X.; Wang, J.; and Zhang, H. 2021 · 2021
Later among the works it cites.
Estimating α \alpha -Rank by Maximizing Information Gain
Rashid, T.; Zhang, C.; and Ciosek, K. 2021 · 2021
Later among the works it cites.
Adversarial dueling bandits
Saha, A.; Koren, T.; and Mansour, Y. 2021 · 2021
Later among the works it cites.
Provably optimal algorithms for generalized linear contextual bandits
Li, L.; Lu, Y.; and Zhou, D. 2017 · 2080
Closest in time.