Fetching the paper…
Reading the bibliography…
We introduce DREAM, a deep reinforcement learning algorithm that finds optimal strategies in imperfect-information games with multiple agents.
Random sampling with a reservoir
Jeffrey S Vitter · 1985
Earlier work this paper cites.
Dynamic programming for partially observable stochastic games
Eric A Hansen, Daniel S Bernstein, and Shlomo Zilberstein · 2004
Earlier work this paper cites.
Bayes’ bluff: Opponent modelling in poker
Finnegan Southey, Michael P Bowling, Bryce Larson, Carmelo Piccione, Neil Burch, Darse Billings, and Chris Rayner · 2005
Earlier work this paper cites.
Regret minimization in games with incomplete information
Martin Zinkevich, Michael Johanson, Michael Bowling, and Carmelo Piccione · 2008
Earlier work this paper cites.
Monte carlo sampling for regret minimization in extensive games
Marc Lanctot, Kevin Waugh, Martin Zinkevich, and Michael Bowling · 2009
Earlier work this paper cites.
A theoretical and empirical analysis of expected sarsa
Harm Van Seijen, Hado Van Hasselt, Shimon Whiteson, and Marco Wiering · 2009
Earlier work this paper cites.
Monte Carlo sampling and regret minimization for equilibrium computation and decision-making in large extensive form games
Marc Lanctot · 2013
Earlier work this paper cites.
Potential-aware imperfect-recall abstraction with earth mover’s distance in imperfect-information games
Sam Ganzfried and Tuomas Sandholm · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Solving large imperfect information games using cfr+
Oskari Tammelin · 2014
Earlier work this paper cites.
Heads-up limit hold’em poker is solved
Michael Bowling, Neil Burch, Michael Johanson, and Oskari Tammelin · 2015
Earlier work this paper cites.
Hierarchical abstraction, distributed equilibrium computation, and post-processing, with application to a champion no-limit texas hold’em agent
Noam Brown, Sam Ganzfried, and Tuomas Sandholm · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2015
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Nando de Freitas, and Marc Lanctot · 2015
Cited alongside, same era.
Solving games with functional regret estimation
Kevin Waugh, Dustin Morrill, James Andrew Bagnell, and Michael Bowling · 2015
Cited alongside, same era.
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver · 2016
Cited alongside, same era.
Equilibrium approximation quality of current no-limit poker bots
Viliam Lisy and Michael Bowling · 2016
The Hanabi challenge: A new frontier for AI research
Nolan Bard, Jakob N Foerster, Sarath Chandar, Neil Burch, Marc Lanctot, H Francis Song, Emilio Parisotto, Vincent Dumoulin, Subhodeep Moitra, Edward Hughes, et al · 2019
Later among the works it cites.
Deep counterfactual regret minimization
Noam Brown, Adam Lerer, Sam Gross, and Tuomas Sandholm · 2019
Later among the works it cites.
Solving imperfect-information games via discounted regret minimization
Noam Brown and Tuomas Sandholm · 2019
Later among the works it cites.
Superhuman AI for multiplayer poker
Noam Brown and Tuomas Sandholm · 2019
Later among the works it cites.
Revisiting cfr+ and alternating updates
Neil Burch, Matej Moravcik, and Martin Schmid · 2019
Later among the works it cites.
Low-variance and zero-variance baselines for extensive-form games
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Noam Brown and Tuomas Sandholm · 2017
Cited alongside, same era.
Regret minimization for partially observable deep reinforcement learning
Peter Jin, Kurt Keutzer, and Sergey Levine · 2017
Cited alongside, same era.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Pérolat, David Silver, and Thore Graepel · 2017
Cited alongside, same era.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisỳ, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Cited alongside, same era.
Trevor Davis, Martin Schmid, and Michael Bowling · 2019
Later among the works it cites.
Rethinking formal models of partially observable multiagent decision making
Vojtěch Kovařík, Martin Schmid, Neil Burch, Michael Bowling, and Viliam Lisỳ · 2019
Later among the works it cites.
Shayegan Omidshafiei, Daniel Hennes, Dustin Morrill, Remi Munos, Julien Perolat, Marc Lanctot, Audrunas Gruslys, Jean-Baptiste Lespiau, and Karl Tuyls · 2019
Later among the works it cites.
Variance reduction in monte carlo counterfactual regret minimization (vr-mccfr) for extensive form games using baselines
Martin Schmid, Neil Burch, Marc Lanctot, Matej Moravcik, Rudolf Kadlec, and Michael Bowling · 2019
Later among the works it cites.
Single deep counterfactual regret minimization
Eric Steinberger · 2019
Later among the works it cites.
Stochastic regret minimization in extensive-form games, 2020
Gabriele Farina, Christian Kroer, and Tuomas Sandholm · 2020
Closest in time.
Double neural counterfactual regret minimization
Hui Li, Kailiang Hu, Shaohua Zhang, Yuan Qi, and Le Song · 2020
Closest in time.
Julien Perolat, Remi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei, Mark Rowland, Pedro Ortega, Neil Burch, Thomas Anthony, David Balduzzi, Bart De Vylder, et al · 2020
Closest in time.