Fetching the paper…
Reading the bibliography…
In this paper, we present exploitability descent, a new algorithm to compute approximate equilibria in two-player zero-sum extensive-form games with imperfect information, by direct policy optimization against worst-case opponents.
Simplified two-person Poker
H. W. Kuhn · 1950
Earlier work this paper cites.
Linear programming and the theory of games
D. Gale, H.W. Kuhn, and A.W. Tucker · 1951
Earlier work this paper cites.
Generalized gradients and applications
Frank H Clarke · 1975
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkadii Semenovich Nemirovsky and David Borisovich Yudin · 1983
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Fast algorithms for finding randomized strategies in game trees
D. Koller, N. Megiddo, and B. von Stengel · 1994
Earlier work this paper cites.
Likelihood ratio gradient estimation for stochastic recursions
Peter W Glynn and Pierre L’ecuyer · 1995
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
S. Hart and A. Mas-Colell · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L Bartlett · 2001
Earlier work this paper cites.
Deep blue
M. Campbell, A. J. Hoane, and F. Hsu · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
M. Zinkevich · 2003
Earlier work this paper cites.
Convex Optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Cited alongside, same era.
Bayes’ bluff: Opponent modelling in poker
Finnegan Southey, Michael Bowling, Bryce Larson, Carmelo Piccione, Neil Burch, Darse Billings, and Chris Rayner · 2005
Cited alongside, same era.
A gradient-based approach for computing Nash equilibria of large sequential games
S. Hoda, A. Gilpin, and J. Pe na · 2007
Cited alongside, same era.
Regret minimization in games with incomplete information
M. Zinkevich, M. Johanson, M. Bowling, and C. Piccione · 2008
Cited alongside, same era.
Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations
Y. Shoham and K. Leyton-Brown · 2009
Cited alongside, same era.
Computer poker: A review
J. Rubin and I. Watson · 2011
Cited alongside, same era.
Solving games with functional regret estimation
Kevin Waugh, Dustin Morrill, J. Andrew Bagnell, and Michael Bowling · 2015
Later among the works it cites.
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Later among the works it cites.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Noam Brown and Tuomas Sandholm · 2017
Later among the works it cites.
Dynamic thresholding and pruning for regret minimization
Noam Brown, Christian Kroer, and Tuomas Sandholm · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Finding optimal abstract strategies in extensive form games
M. Johanson, N. Bard, N. Burch, and M. Bowling · 2012
Cited alongside, same era.
Online learning and online convex optimization
Shai Shalev-Shwartz et al · 2012
Cited alongside, same era.
Monte Carlo Sampling and Regret Minimization for Equilibrium Computation and Decision-Making in Large Extensive Form Games
Marc Lanctot · 2013
Cited alongside, same era.
A unified view of large-scale zero-sum equilibrium computation
Kevin Waugh and J. Andrew Bagnell · 2014
Cited alongside, same era.
Heads-up Limit Hold’em Poker is solved
Michael Bowling, Neil Burch, Michael Johanson, and Oskari Tammelin · 2015
Cited alongside, same era.
Introduction to online convex optimization
Elad Hazan · 2015
Cited alongside, same era.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Perolat, David Silver, and Thore Graepel · 2017
Later among the works it cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisý, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling · 2017
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Later among the works it cites.
Actor-critic policy optimization in partially observable multiagent environments
Sriram Srinivasan, Marc Lanctot, Vinicius Zambaldi, Julien Pérolat, Karl Tuyls, Rémi Munos, and Michael Bowling · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
R. Sutton and A. Barto · 2018
Later among the works it cites.
Equilibrium finding via asymmetric self-play reinforcement learning
Jie Tang, Keiran Paster, and Pieter Abbeel · 2018
Later among the works it cites.
Deep counterfactual regret minimization
Noam Brown, Adam Lerer, Sam Gross, and Tuomas Sandholm · 2019
Closest in time.
Computing approximate equilibria in sequential adversarial games by exploitability descent
Edward Lockhart, Marc Lanctot, Julien Pérolat, Jean-Baptiste Lespiau, Dustin Morrill, Finbarr Timbers, and Karl Tuyls · 2019
Closest in time.