Fetching the paper…
Reading the bibliography…
Policy-based methods with function approximation are widely used for solving two-player zero-sum games with large state and/or action spaces.
Iterative solution of games by fictitious play
George W Brown · 1951
Earlier work this paper cites.
An iterative method of solving a game
Julia Robinson · 1951
Earlier work this paper cites.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Stochastic and shortest path games: theory and algorithms
Stephen David Patek · 1997
Earlier work this paper cites.
A tensorial approach to computational continuum mechanics using object-oriented techniques
Henry G Weller, Gavin Tabor, Hrvoje Jasak, and Christer Fureby · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2002
Earlier work this paper cites.
Basis function adaptation in temporal difference reinforcement learning
Ishai Menache, Shie Mannor, and Nahum Shimkin · 2005
Earlier work this paper cites.
Error bounds for approximate value iteration
Rémi Munos · 2005
Earlier work this paper cites.
Progress in approximate Nash equilibria
Constantinos Daskalakis, Aranyak Mehta, and Christos Papadimitriou · 2007
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Regret minimization in games with incomplete information
Martin Zinkevich, Michael Johanson, Michael Bowling, and Carmelo Piccione · 2008
Earlier work this paper cites.
Reinforcement learning for mapping instructions to actions
S. R. K. Branavan, Harr Chen, Luke S. Zettlemoyer, and Regina Barzilay · 2009
Earlier work this paper cites.
Softmax-margin crfs: Training log-linear models with cost functions
Kevin Gimpel and Noah A Smith · 2010
Earlier work this paper cites.
Competitive Markov decision processes
Jerzy Filar and Koos Vrieze · 2012
Earlier work this paper cites.
Approximate modified policy iteration
Bruno Scherrer, Victor Gabillon, Mohammad Ghavamzadeh, and Matthieu Geist · 2012
Earlier work this paper cites.
Actor-critic reinforcement learning with energy-based policies
Nicolas Heess, David Silver, and Yee Whye Teh · 2013
Earlier work this paper cites.
Optimization, learning, and games with predictable sequences
Sasha Rakhlin and Karthik Sridharan · 2013
Earlier work this paper cites.
Approximate policy iteration schemes: a comparison
Bruno Scherrer · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Fictitious self-play in extensive-form games
Johannes Heinrich, Marc Lanctot, and David Silver · 2015
Cited alongside, same era.
Abstraction selection in model-based reinforcement learning
Nan Jiang, Alex Kulesza, and Satinder Singh · 2015
Cited alongside, same era.
Approximate dynamic programming for two-player zero-sum Markov games
Julien Perolat, Bruno Scherrer, Bilal Piot, and Olivier Pietquin · 2015
Cited alongside, same era.
Approximate modified policy iteration and its application to the game of tetris
Bruno Scherrer, Mohammad Ghavamzadeh, Victor Gabillon, Boris Lesner, and Matthieu Geist · 2015
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 2019
Later among the works it cites.
Solving imperfect-information games via discounted regret minimization
Noam Brown and Tuomas Sandholm · 2019
Later among the works it cites.
Global convergence of policy gradient for sequential zero-sum linear quadratic dynamic games
Jingjing Bu, Lillian J Ratliff, and Mehran Mesbahi · 2019
Later among the works it cites.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Later among the works it cites.
Computing approximate equilibria in sequential adversarial games by exploitability descent
Edward Lockhart, Marc Lanctot, Julien Pérolat, Jean-Baptiste Lespiau, Dustin Morrill, Finbarr Timbers, and Karl Tuyls · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Deep learning for reward design to improve monte carlo tree search in atari games
Xiaoxiao Guo, Satinder Singh, Richard Lewis, and Honglak Lee · 2016
Cited alongside, same era.
On the use of non-stationary strategies for solving two-player zero-sum Markov games
Julien Pérolat, Bilal Piot, Bruno Scherrer, and Olivier Pietquin · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Computing approximate Nash equilibria in polymatrix games
Argyrios Deligkas, John Fearnley, Rahul Savani, and Paul Spirakis · 2017
Cited alongside, same era.
Learning with opponent-learning awareness
Jakob N Foerster, Richard Y Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Later among the works it cites.
Optimistic mirror descent in saddle-point problems: Going the extra(-gradient) mile
Panayotis Mertikopoulos, Bruno Lecouat, Houssam Zenati, Chuan-Sheng Foo, Vijay Chandrasekhar, and Georgios Piliouras · 2019
Later among the works it cites.
Elf opengo: An analysis and open reimplementation of alphazero
Yuandong Tian, Jerry Ma, Qucheng Gong, Shubho Sengupta, Zhuoyuan Chen, James Pinkerton, and Larry Zitnick · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
Provable q-iteration with l infinity guarantees and function approximation
Ming Yu, Zhuoran Yang, Mengdi Wang, and Zhaoran Wang · 2019
Later among the works it cites.
Policy optimization provably converges to Nash equilibria in zero-sum linear quadratic games
Kaiqing Zhang, Zhuoran Yang, and Tamer Basar · 2019
Later among the works it cites.
Optimality and approximation with policy gradient methods in Markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2020
Later among the works it cites.
Communication complexity of approximate Nash equilibria
Yakov Babichenko and Aviad Rubinstein · 2020
Later among the works it cites.
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Later among the works it cites.
Near-optimal reinforcement learning with self-play
Yu Bai, Chi Jin, and Tiancheng Yu · 2020
Later among the works it cites.
Fast global convergence of natural policy gradient methods with entropy regularization
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2020
Later among the works it cites.
Independent policy gradient methods for competitive reinforcement learning
Constantinos Daskalakis, Dylan J Foster, and Noah Golowich · 2020
Later among the works it cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized MDPs
Lior Shani, Yonathan Efroni, and Shie Mannor · 2020
Later among the works it cites.
Global convergence of policy gradient methods to (almost) locally optimal policies
Kaiqing Zhang, Alec Koppel, Hao Zhu, and Tamer Basar · 2020
Later among the works it cites.
From poincaré recurrence to convergence in imperfect information games: Finding equilibrium via regularization
Julien Perolat, Remi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei, Mark Rowland, Pedro Ortega, Neil Burch, Thomas Anthony, David Balduzzi, Bart De Vylder, et al · 2021
Closest in time.