Fetching the paper…
Reading the bibliography…
Hannan consistency, or no external regret, is a~key concept for learning in games.
Approximation to bayes risk in repeated play
James Hannan · 1957
Earlier work this paper cites.
Goofspiel — the game of pure strategy
Sheldon M Ross · 1971
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Gambling in a rigged casino: The adversarial multi-armed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 1995
Earlier work this paper cites.
The theory of learning in games
Drew Fudenberg, David K Levine, et al · 1998
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
Sergiu Hart and Andreu Mas-Colell · 2000
Earlier work this paper cites.
A reinforcement procedure leading to correlated equilibrium
Sergiu Hart and Andreu Mas-Colell · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Multiagent learning using a variable learning rate
Michael Bowling and Manuela Veloso · 2002
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicoló Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 2003
Earlier work this paper cites.
Solving the Oshi-zumo game
Michael Buro · 2004
Earlier work this paper cites.
Online convex optimization in the bandit setting: gradient descent without a gradient
Abraham D Flaxman, Adam Tauman Kalai, and H Brendan McMahan · 2005
Earlier work this paper cites.
Incomplete information and internal regret in prediction of individual sequences
Gilles Stoltz · 2005
Earlier work this paper cites.
Prediction, Learning, and Games
Nicolo Cesa-Bianchi and Gabor Lugosi · 2006
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Learning, regret minimization, and equilibria
A. Blum and Y. Mansour · 2007
Cited alongside, same era.
Bandit algorithms for tree search
Pierre-Arnuad Coquelin and Remi Munos · 2007
Cited alongside, same era.
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom · 2007
Cited alongside, same era.
Multi-armed bandit problems with dependent arms
Sandeep Pandey, Deepayan Chakrabarti, and Deepak Agarwal · 2007
Cited alongside, same era.
Regret minimization in games with incomplete information
Martin Zinkevich, Michael Johanson, Michael Bowling, and Carmelo Piccione · 2007
Cited alongside, same era.
Simulation-based approach to general game playing
Hilmar Finnsson and Yngvi Björnsson · 2008
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Later among the works it cites.
Generalized sampling and variance in counterfactual regret minimization
Richard G Gibson, Marc Lanctot, Neil Burch, Duane Szafron, and Michael Bowling · 2012
Later among the works it cites.
Comparison of different selection strategies in monte-carlo tree search for the game of Tron
Pierre Perick, David L. St-Pierre, Francis Maes, and Damien Ernst · 2012
Later among the works it cites.
Computer solution to the game of pure strategy
Glenn C. Rhoads and Laurent Bartholdi · 2012
Later among the works it cites.
Alpha-beta pruning for games with simultaneous moves
Abdallah Saffidine, Hilmar Finnsson, and Michael Buro · 2012
Later among the works it cites.
Online learning in markov decision processes with adversarially chosen transition probability distributions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Monte carlo sampling for regret minimization in extensive games
Marc Lanctot, Kevin Waugh, Martin Zinkevich, and Michael Bowling · 2009
Cited alongside, same era.
Comparing UCT versus CFR in simultaneous games
Mohammad Shafiei, Nathan Sturtevant, and Jonathan Schaeffer · 2009
Cited alongside, same era.
Abstraction in large extensive games
Kevin Waugh · 2009
Cited alongside, same era.
Arbitrarily modulated markov decision processes
Jia Yuan Yu and Shie Mannor · 2009
Cited alongside, same era.
The online loop-free stochastic shortest-path problem
Gergely Neu, András György, and Csaba Szepesvári · 2010
Cited alongside, same era.
Monte-carlo tree search and rapid action value estimation in computer go
Sylvain Gelly and David Silver · 2011
Cited alongside, same era.
Yasin Abbasi, Peter L Bartlett, Varun Kanade, Yevgeny Seldin, and Csaba Szepesvári · 2013
Later among the works it cites.
Using Double-oracle Method and Serialized Alpha-Beta Search for Pruning in Simultaneous Move Games
Branislav Bošanský, Viliam Lisý, Jiří Čermák, Roman Vítek, and Michal Pěchouček · 2013
Later among the works it cites.
Convergence of monte carlo tree search in simultaneous move games
Viliam Lisý, Vojta Kovařík, Marc Lanctot, and Branislav Bošanský · 2013
Later among the works it cites.
Monte Carlo tree search in simultaneous move games with applications to Goofspiel
Marc Lanctot, Viliam Lisý, and Mark HM Winands · 2014
Later among the works it cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Later among the works it cites.
Monte Carlo tree search variants for simultaneous move games
M. J. W. Tak, M. H. M. Winands, and M. Lanctot · 2014
Later among the works it cites.
Online monte carlo counterfactual regret minimization for search in imperfect information games
Viliam Lisy, Marc Lanctot, and Michael Bowling · 2015
Later among the works it cites.
Algorithms for computing strategies in two-player simultaneous move games
Branislav Bošanský, Viliam Lisý, Marc Lanctot, Jiří Čermák, and Mark HM Winands · 2016
Later among the works it cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisý, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling · 2017
Later among the works it cites.
Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
Noam Brown and Tuomas Sandholm · 2018
Closest in time.