Fetching the paper…
Reading the bibliography…
We investigate the increasingly important and common game-solving setting where we do not have an explicit description of the game but only oracle access to it through gameplay, such as in financial or military simulations and computer games.
Zur theorie der gesellschaftsspiele
John von Neumann · 1928
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R. Thompson · 1933
Earlier work this paper cites.
An invariant form for the prior probability in estimation problems
Harold Jeffreys · 1946
Earlier work this paper cites.
Stochastic games
Lloyd S. Shapley · 1953
Earlier work this paper cites.
On nonterminating stochastic games
Alan J. Hoffman and Richard M. Karp · 1966
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L. Littman · 1994
Earlier work this paper cites.
Multiagent reinforcement learning in the iterated prisoner’s dilemma
Tuomas W. Sandholm and Robert H. Crites · 1996
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Bayesian q-learning
Richard Dearden, Nir Friedman, and Stuart Russell · 1998
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Reinforcement learning to play an optimal nash equilibrium in team markov games
Xiaofeng Wang and Tuomas W. Sandholm · 2002
Earlier work this paper cites.
Nash q-learning for general-sum stochastic games
Junling Hu and Michael P. Wellman · 2003
Cited alongside, same era.
Pure exploration in multi-armed bandits problems
Sébastien Bubeck, Rémi Munos, and Gilles Stoltz · 2009
Cited alongside, same era.
Computing equilibria in multiplayer stochastic games of imperfect information
Sam Ganzfried and Tuomas W. Sandholm · 2009
Cited alongside, same era.
Strategic analysis with simulation-based games
Yevgeniy Vorobeychik and Michael P. Wellman · 2009
Cited alongside, same era.
A minimum relative entropy principle for learning and acting
Pedro A. Ortega and Daniel A. Braun · 2010
Cited alongside, same era.
On bayesian upper confidence bounds for bandit problems
Emilie Kaufmann, Olivier Cappe, and Aurelien Garivier · 2012
Cited alongside, same era.
Ucb exploration via q-ensembles, 2017
Richard Y. Chen, Szymon Sidor, Pieter Abbeel, and John Schulman · 2017
Later among the works it cites.
Why is posterior sampling better than optimism for reinforcement learning?
Ian Osband and Benjamin Van Roy · 2017
Later among the works it cites.
The uncertainty bellman equation and exploration
Brendan O’Donoghue, Ian Osband, Rémi Munos, and Volodymyr Mnih · 2018
Later among the works it cites.
A tutorial on thompson sampling
Daniel J. Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2018
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning, 2019
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Dębiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, Rafal Józefowicz, Scott Gray, Catherine Olsson, Jakub Pachocki, Michael Petrov, Henrique Pondé de Oliveira Pinto, Jonathan Raiman, Tim Salimans, Jeremy Schlatter, Jonas Schneider, Szymon Sidor, Ilya Sutskever, Jie Tang, Filip Wolski, and Susan Zhang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to optimize via posterior sampling
Daniel Russo and Benjamin Van Roy · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Maximin action identification: a new bandit framework for games
Aurélien Garivier, Emilie Kaufmann, and Wouter M. Koolen · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Marc G. Bellemare, Will Dabney, and Rémi Munos · 2017
Cited alongside, same era.
Deep q-learning for nash equilibria: Nash-dqn, 2019
Philippe Casgrain, Brian Ning, and Sebastian Jaimungal · 2019
Later among the works it cites.
Learning probably approximately correct maximin strategies in simulation-based games with infinite strategy spaces, 2019
Alberto Marchesi, Francesco Trovò, and Nicola Gatti · 2019
Later among the works it cites.
Distributional reinforcement learning for efficient exploration
Borislav Mavrin, Hengshuai Yao, Linglong Kong, Kaiwen Wu, and Yaoliang Yu · 2019
Later among the works it cites.
Deep exploration via randomized value functions
Ian Osband, Benjamin Van Roy, Daniel J. Russo, and Zheng Wen · 2019
Later among the works it cites.
Learning deviation payoffs in simulation-based games
Samuel Sokota, Caleb Ho, and Bryce Wiedenbeck · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M. Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H. Choi, Richard Powell, Timo Ewalds, Petko Georgiev, Junhyuk Oh, Dan Horgan, Manuel Kroiss, Ivo Danihelka, Aja Huang, Laurent Sifre, Trevor Cai, John P. Agapiou, Max Jaderberg, Alexander S. Vezhnevets, Rémi Leblond, Tobias Pohlen, Valentin Dalibard, David Budden, Yury Sulsky, James Molloy, Tom L. Paine, Caglar Gulcehre, Ziyu Wang, Tobias Pfaff, Yuhuai Wu, Roman Ring, Dani Yogatama, Dario Wünsch, Katrina McKinney, Oliver Smith, Tom Schaul, Timothy Lillicrap, Koray Kavukcuoglu, Demis Hassabis, Chris Apps, and David Silver · 2019
Later among the works it cites.