Fetching the paper…
Reading the bibliography…
Deep reinforcement learning (RL) has achieved outstanding results in recent years, which has led a dramatic increase in the number of methods and applications.
On the theory of brownian motion
G.E.Uhlenbeck and L.S.Ornstein · 1930
Earlier work this paper cites.
Equilibrium points in n-person games
John F. Nash · 1950
Earlier work this paper cites.
An iterative method of solving a game
Julia Robinson · 1951
Earlier work this paper cites.
Stochastic games
L. S. Shapley · 1953
Earlier work this paper cites.
Extensive games and the problem of information
H.W. Kuhn · 1953
Earlier work this paper cites.
On the convergence of learning processes in a 2 x 2 non-zero sum game, 1961
Koichi Miyasawa · 1961
Earlier work this paper cites.
Equilibrium in a stochastic
Arlington M Fink et al · 1964
Earlier work this paper cites.
Subjectivity and correlation in randomized strategies
Robert Aumann · 1974
Earlier work this paper cites.
Isolated invariant sets and the Morse index
Charles Conley · 1978
Earlier work this paper cites.
The multi-armed bandit problem: Decomposition and computation
Michael N. Katehakis and Arthur F. Veinott · 1987
Earlier work this paper cites.
Learning from Delayed Rewards
C. J. C. H. Watkins · 1989
Earlier work this paper cites.
Optimality and Equilibria in Stochastic Games
F Thuijsman · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Learning mixed equilibria
Drew Fudenberg and David M. Kreps · 1993
Earlier work this paper cites.
Asynchronous stochastic approximation and q-learning
John N. Tsitsiklis · 1994
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L. Littman · 1994
Earlier work this paper cites.
Potential games
Dov Monderer and Lloyd S. Shapley · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The Theory of Learning in Games
Drew Fudenberg and David K. Levine · 1998
Earlier work this paper cites.
The convergence of fictitious play in 3 x 3 games with strategic complementarities
Sunku Hahn · 1999
Earlier work this paper cites.
Multiagent systems: A survey from a machine learning perspective
Peter Stone and Manuela Veloso · 2000
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
Sergiu Hart and Andreu Mas-Colell · 2000
Earlier work this paper cites.
Actor-critic algorithms
Vijay R. Konda and John N. Tsitsiklis · 2000
Earlier work this paper cites.
Convergence problems of general-sum multiagent reinforcement learning
Michael Bowling · 2000
Earlier work this paper cites.
From nash and brown to maynard smith: Equilibria, dynamics and ess
J. Hofbauer · 2000
Earlier work this paper cites.
A weakened form of fictitious play in two-person zero-sum games
Ben Van Der Genugten · 2000
Earlier work this paper cites.
Value-function reinforcement learning in markov games
Michael L. Littman · 2001
Earlier work this paper cites.
Friend-or-foe q-learning in general-sum games
Michael L. Littman · 2001
Earlier work this paper cites.
A survey and critique of multiagent deep reinforcement learning
Pablo Hernandez-Leal, Bilal Kartal, and Matthew E. Taylor · 2001
Earlier work this paper cites.
On the global convergence of stochastic fictitious play
Josef Hofbauer and William H. Sandholm · 2002
Earlier work this paper cites.
Analyzing complex strategic interactions in multi-agent systems
William E. Walsh, Rajarshi Das, Gerald Tesauro, and Jeffrey O. Kephart · 2002
Earlier work this paper cites.
Nash q-learning for general-sum stochastic games
Junling Hu and Michael P. Wellman · 2003
Earlier work this paper cites.
Correlated-q learning
Amy Greenwald and Keith Hall · 2003
Earlier work this paper cites.
Asymmetric multiagent reinforcement learning
Ville Könönen · 2003
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
H. Brendan McMahan, Geoffery J Gordon, and Avrim Blum · 2003
Earlier work this paper cites.
Cyclic equilibria in markov games
Martin Zinkevich, Amy R. Greenwald, and Michael L. Littman · 2005
Cited alongside, same era.
Brown’s original fictitious play
Ulrich Berger · 2005
Cited alongside, same era.
Generalised weakened fictitious play
David S. Leslie and E.J. Collins · 2006
Cited alongside, same era.
Bounds for regret-matching algorithms, 2006
Amy Greenwald, Zheng Li, and Casey Marks · 2006
Cited alongside, same era.
Settling the complexity of 2-player nash-equilibrium
Xi Chen, Xiaotie Deng, and Shang-Hua Teng · 2007
Cited alongside, same era.
Regret minimization in games with incomplete information
Martin Zinkevich, Michael Johanson, and Michael Bowlingand Carmelo Piccione · 2007
Cited alongside, same era.
A survey of learning in multiagent environments: Dealing with non-stationarity, 2017
Tim Baarslag Enrique Munoz de Cote Pablo Hernandez-Leal, Michael Kaisers · 2017
Later among the works it cites.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Can deep reinforcement learning solve erdos-selfridge-spencer games?, 2017
Maithra Raghu, Alex Irpan, Jacob Andreas, Robert Kleinberg, Quoc V. Le, and Jon Kleinberg · 2017
Later among the works it cites.
Multi-agent reinforcement learning in sequential social dilemmas
Joel Z. Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, and Thore Graepel · 2017
Later among the works it cites.
Emergent complexity via multi-agent competition
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch · 2017
Later among the works it cites.
Learning nash equilibrium for general-sum markov games from batch data
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yoav Shoham, Rob Powers, and Trond Grenager · 2007
Cited alongside, same era.
Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations
Yoav Shoham and Kevin Leyton-Brown · 2009
Cited alongside, same era.
Monte carlo sampling for regret minimization in extensive games
Marc Lanctot, Kevin Waugh, Martin Zinkevich, and Michael Bowling · 2009
Cited alongside, same era.
Using counterfactual regret minimization to create competitive multiplayer poker agents
Nick Abou Risk and Duane Szafron · 2010
Cited alongside, same era.
The world of independent learners is not markovian, 2011
Guillaume J. Laurent, Laetitia Matignon, and Nadine Le-Fort Piat · 2011
Cited alongside, same era.
A double oracle algorithm for zero-sum security games on graphs
Manish Jain, Dmytro Korzhyk, Ondˇrej Vanek, Vincent Conitzer, Michal Pechoucek, and Milind Tambe · 2011
Cited alongside, same era.
Julien Pérolat, Florian Strub, Bilal Piot, and Olivier Pietquin · 2017
Later among the works it cites.
Neural fictitious self-play in imperfect information games with many players
Keigo Kawamura, Naoki Mizukami, and Yoshimasa Tsuruoka · 2017
Later among the works it cites.
Learning with opponent-learning awareness
Jakob N. Foerster, Richard Y. Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch · 2017
Later among the works it cites.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Perolat, David Silver, and Thore Graepel · 2017
Later among the works it cites.
Regret minimization for partially observable deep reinforcement learning
Peter Jin, Kurt Keutzer, and Sergey Levine · 2017
Later among the works it cites.
Libratus: The superhuman ai for no-limit poker
Noam Brown and Tuomas Sandholm · 2017
Later among the works it cites.
Openai five
OpenAI · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Later among the works it cites.
Actor-critic fictitious play in simultaneous move multistage games
Julien Perolat, Bilal Piot, and Olivier Pietquin · 2018
Later among the works it cites.
Adversarial reinforcement learning for observer design in autonomous systems under cyber attacks, 2018
Abhishek Gupta and Zhaoyuan Yang · 2018
Later among the works it cites.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sainbayar Sukhbaatar, Zeming Lin, Ilya Kostrikov, Gabriel Synnaeve, Arthur Szlam, and Rob Fergus · 2018
Later among the works it cites.
Deep counterfactual regret minimization
Noam Brown, Adam Lerer, Sam Gross, and Tuomas Sandholm · 2018
Later among the works it cites.
Solving imperfect-information games via discounted regret minimization, 2018
Noam Brown and Tuomas Sandholm · 2018
Later among the works it cites.
Finding friend and foe in multi-agent games, 2019
Jack Serrino, Max Kleiman-Weiner, David C. Parkes, and Joshua B. Tenenbaum · 2019
Later among the works it cites.
α \alpha -rank: Multi-agent evaluation by evolution, 2019
Shayegan Omidshafiei, Christos Papadimitriou, Georgios Piliouras, Karl Tuyls, Mark Rowland, Jean-Baptiste Lespiau, Wojciech M. Czarnecki, Marc Lanctot, Julien Perolat, and Remi Munos · 2019
Later among the works it cites.
Multi-agent adversarial inverse reinforcement learning
Lantao Yu, Jiaming Song, and Stefano Ermon · 2019
Later among the works it cites.
A generalized training approach for multiagent learning
Paul Muller, Shayegan Omidshafiei, Mark Rowland, Karl Tuyls, Julien Perolat, Siqi Liu, Daniel Hennes, Luke Marris, Marc Lanctot, Edward Hughes, Zhe Wang, Guy Lever, Nicolas Heess, Thore Graepel, and Remi Munos · 2019
Later among the works it cites.
α α \alpha^{\alpha} -rank: Practically scaling
Yaodong Yang, Rasul Tutunov, Phu Sakulwongtana, and Haitham Bou Ammar · 2019
Later among the works it cites.
Computing stackelberg equilibria of large general-sum games, 2019
Avrim Blum, Nika Hagtalab, MohammadTaghi Hajiaghayi, and Saeed Seddighin · 2019
Later among the works it cites.
Multiagent evaluation under incomplete information
Mark Rowland, Shayegan Omidshafiei, Karl Tuyls, Julien Perolat, Michal Valko, Georgios Piliouras, and Remi Munos · 2019
Later among the works it cites.
Monte carlo neural fictitious self-play: Approach to approximate nash equilibrium of imperfect-information games, 2019
Li Zhang, Wei Wang, Shijian Li, and Gang Pan · 2019
Later among the works it cites.
Neural fictitious self-play on elf mini-rts, 2019
Keigo Kawamura and Yoshimasa Tsuruoka · 2019
Later among the works it cites.
Bounds for approximate regret-matching algorithms, 2019
Ryan D’Orazio, Dustin Morrill, and James R. Wright · 2019
Later among the works it cites.
Alternative function approximation parameterizations for solving games: An analysis of
Ryan D’Orazio, Dustin Morrill, James R. Wright, and Michael Bowling · 2019
Later among the works it cites.
Stable-predictive optimistic counterfactual regret minimization
Gabriele Farina, Christian Kroer, Noam Brown, and Tuomas Sandholm · 2019
Later among the works it cites.
Variance reduction in monte carlo counterfactual regret minimization (vr-mccfr) for extensive form games using baselines
Martin Schmid, Neil Burch, Marc Lanctot, Matej Moravcik, Rudolf Kadlec, and Michael Bowling · 2019
Later among the works it cites.
Single deep counterfactual regret minimization, 2019
Eric Steinberger · 2019
Later among the works it cites.
Double neural counterfactual regret minimization
Hui Li, Kailiang Hu, Zhibang Ge, Tao Jiang, Yuan Qi, and Le Song · 2019
Later among the works it cites.
Combining no-regret and q-learning, 2019
Katja Hofmann Ian A. Kash, Michael Sullins · 2019
Later among the works it cites.