Fetching the paper…
Reading the bibliography…
The combination of deep reinforcement learning and search at both training and test time is a powerful paradigm that has led to a number of successes in single-agent settings and perfect-information games, best exemplified by AlphaZero.
Programming a computer for playing chess
Claude E Shannon · 1950
Earlier work this paper cites.
Iterative solution of games by fictitious play
George W Brown · 1951
Earlier work this paper cites.
Non-cooperative games
John Nash · 1951
Earlier work this paper cites.
Some studies in machine learning using the game of checkers
Arthur L Samuel · 1959
Earlier work this paper cites.
Convex analysis
R Tyrrell Rockafellar · 1970
Earlier work this paper cites.
Agreeing to disagree
Robert J Aumann · 1976
Earlier work this paper cites.
TD-Gammon, a self-teaching backgammon program, achieves master-level play
Gerald Tesauro · 1994
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra · 1998
Earlier work this paper cites.
A weakened form of fictitious play in two-person zero-sum games
Ben Van der Genugten · 2000
Earlier work this paper cites.
Dynamic programming for partially observable stochastic games
Eric A Hansen, Daniel S Bernstein, and Shlomo Zilberstein · 2004
Earlier work this paper cites.
Optimal rhode island hold’em poker
Andrew Gilpin and Tuomas Sandholm · 2005
Earlier work this paper cites.
A competitive texas hold’em poker player via automated abstraction and real-time equilibrium computation
Andrew Gilpin and Tuomas Sandholm · 2006
Earlier work this paper cites.
Generalised weakened fictitious play
David S Leslie and Edmund J Collins · 2006
Earlier work this paper cites.
Combining online and offline knowledge in uct
Sylvain Gelly and David Silver · 2007
Earlier work this paper cites.
Regret minimization in games with incomplete information
Martin Zinkevich, Michael Johanson, Michael Bowling, and Carmelo Piccione · 2008
Earlier work this paper cites.
Smoothing techniques for computing nash equilibria of sequential games
Samid Hoda, Andrew Gilpin, Javier Pena, and Tuomas Sandholm · 2010
Earlier work this paper cites.
Online optimization with gradual variations
Chao-Kai Chiang, Tianbao Yang, Chia-Jung Lee, Mehrdad Mahdavi, Chi-Jen Lu, Rong Jin, and Shenghuo Zhu · 2012
Earlier work this paper cites.
Finding optimal abstract strategies in extensive-form games
Michael Johanson, Nolan Bard, Neil Burch, and Michael Bowling · 2012
Earlier work this paper cites.
Decentralized stochastic control with partial history sharing: A common information approach
Ashutosh Nayyar, Aditya Mahajan, and Demosthenis Teneketzis · 2013
Earlier work this paper cites.
Sufficient plan-time statistics for decentralized pomdps
Frans Adriaan Oliehoek · 2013
Earlier work this paper cites.
Online learning with predictable sequences
Alexander Rakhlin and Karthik Sridharan · 2013
Earlier work this paper cites.
Solving imperfect information games using decomposition
Neil Burch, Michael Johanson, and Michael Bowling · 2014
Cited alongside, same era.
Potential-aware imperfect-recall abstraction with earth mover’s distance in imperfect-information games
Sam Ganzfried and Tuomas Sandholm · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Heads-up limit hold’em poker is solved
Michael Bowling, Neil Burch, Michael Johanson, and Oskari Tammelin · 2015
Cited alongside, same era.
Hierarchical abstraction, distributed equilibrium computation, and post-processing, with application to a champion no-limit texas hold’em agent
Noam Brown, Sam Ganzfried, and Tuomas Sandholm · 2015
Cited alongside, same era.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisỳ, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
Depth-limited solving for imperfect-information games
Noam Brown, Tuomas Sandholm, and Brandon Amos · 2018
Later among the works it cites.
Aivat: A new variance reduction technique for agent evaluation in imperfect information games
Neil Burch, Martin Schmid, Matej Moravcik, Dustin Morill, and Michael Bowling · 2018
Later among the works it cites.
Solving large sequential games with the excessive gap technique
Christian Kroer, Gabriele Farina, and Tuomas Sandholm · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Simultaneous abstraction and equilibrium finding in games
Noam Brown and Tuomas Sandholm · 2015
Cited alongside, same era.
Endgame solving in large imperfect-information games
Sam Ganzfried and Tuomas Sandholm · 2015
Cited alongside, same era.
Fast convergence of regularized learning in games
Vasilis Syrgkanis, Alekh Agarwal, Haipeng Luo, and Robert E Schapire · 2015
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Cited alongside, same era.
Baby tartanian8: Winning agent from the 2016 annual computer poker competition
Noam Brown and Tuomas Sandholm · 2016
Cited alongside, same era.
Strategy-based warm starting for regret minimization in games
Noam Brown and Tuomas Sandholm · 2016
Cited alongside, same era.
Optimally solving dec-pomdps as continuous-state mdps
Jilles Steeve Dibangoye, Christopher Amato, Olivier Buffet, and François Charpillet · 2016
Cited alongside, same era.
Faster algorithms for extensive-form game solving via improved smoothing functions
Christian Kroer, Kevin Waugh, Fatma Kılınç-Karzan, and Tuomas Sandholm · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Later among the works it cites.
Deep counterfactual regret minimization
Noam Brown, Adam Lerer, Sam Gross, and Tuomas Sandholm · 2019
Later among the works it cites.
Solving imperfect-information games via discounted regret minimization
Noam Brown and Tuomas Sandholm · 2019
Later among the works it cites.
Superhuman AI for multiplayer poker
Noam Brown and Tuomas Sandholm · 2019
Later among the works it cites.
Bayesian action decoder for deep multi-agent reinforcement learning
Jakob Foerster, Francis Song, Edward Hughes, Neil Burch, Iain Dunning, Shimon Whiteson, Matthew Botvinick, and Michael Bowling · 2019
Later among the works it cites.
Solving partially observable stochastic games with public observations
Karel Horák and Branislav Bošanskỳ · 2019
Later among the works it cites.
Problems with the efg formalism: a solution attempt using observations
Vojtěch Kovařík and Viliam Lisỳ · 2019
Later among the works it cites.
Rethinking formal models of partially observable multiagent decision making
Vojtěch Kovařík, Martin Schmid, Neil Burch, Michael Bowling, and Viliam Lisỳ · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2019
Later among the works it cites.
Value functions for depth-limited solving in imperfect-information games beyond poker
Dominik Seitz, Vojtech Kovarík, Viliam Lisỳ, Jan Rudolf, Shuo Sun, and Karel Ha · 2019
Later among the works it cites.
Finding friend and foe in multi-agent games
Jack Serrino, Max Kleiman-Weiner, David C Parkes, and Josh Tenenbaum · 2019
Later among the works it cites.
Monte carlo continual resolving for online strategy computation in imperfect information games
Michal Šustr, Vojtěch Kovařík, and Viliam Lisỳ · 2019
Later among the works it cites.
Improving policies via search in cooperative partially observable games
Adam Lerer, Hengyuan Hu, Jakob Foerster, and Noam Brown · 2020
Closest in time.