Fetching the paper…
Reading the bibliography…
Trust region methods are widely applied in single-agent reinforcement learning problems due to their monotonic performance-improvement guarantee at every iteration.
Equilibrium points in n-person games
John F Nash · 1950
Earlier work this paper cites.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
The theory of learning in games
Drew Fudenberg, Fudenberg Drew, David K Levine, and David K Levine · 1998
Earlier work this paper cites.
Nash convergence of gradient dynamics in general-sum games
Satinder Singh, Michael Kearns, and Yishay Mansour · 2000
Earlier work this paper cites.
An efficient, exact algorithm for solving tree-structured graphical games
Michael L. Littman, Michael J. Kearns, and Satinder P. Singh · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham M. Kakade and John Langford · 2002
Earlier work this paper cites.
Multiagent learning using a variable learning rate
Michael Bowling and Manuela Veloso · 2002
Earlier work this paper cites.
Reducing the time complexity of the derandomized evolution strategy with covariance matrix adaptation (cma-es)
Nikolaus Hansen, Sibylle D Müller, and Petros Koumoutsakos · 2003
Earlier work this paper cites.
Playing large games using simple strategies
Richard J Lipton, Evangelos Markakis, and Aranyak Mehta · 2003
Earlier work this paper cites.
Extragradient approach to the solution of two person non-zero sum games
Anatoly Antipin · 2003
Earlier work this paper cites.
Convergent multiple-timescales reinforcement learning algorithms in normal form games
David S Leslie, EJ Collins, et al · 2003
Earlier work this paper cites.
On nash equilibria in stochastic games
Krishnendu Chatterjee, Rupak Majumdar, and Marcin Jurdziński · 2004
Earlier work this paper cites.
Existence of multiagent equilibria with limited agents
Michael Bowling and Manuela Veloso · 2004
Earlier work this paper cites.
The complexity of pure nash equilibria
Alex Fabrikant, Christos Papadimitriou, and Kunal Talwar · 2004
Earlier work this paper cites.
Individual q-learning in normal form games
David S Leslie and Edmund J Collins · 2005
Earlier work this paper cites.
Methods for empirical game-theoretic analysis
Michael P Wellman · 2006
Earlier work this paper cites.
Computing nash equilibria: Approximation and smoothed complexity
Xi Chen, Xiaotie Deng, and Shang-Hua Teng · 2006
Earlier work this paper cites.
Multiagent systems: Algorithmic, game-theoretic, and logical foundations
Yoav Shoham and Kevin Leyton-Brown · 2008
Earlier work this paper cites.
The complexity of computing a nash equilibrium
Constantinos Daskalakis, Paul W Goldberg, and Christos H Papadimitriou · 2009
Earlier work this paper cites.
Generalization risk minimization in empirical game models
Patrick R Jordan and Michael P Wellman · 2009
Earlier work this paper cites.
Multi-agent reinforcement learning: An overview
Lucian Buşoniu, Robert Babuška, and Bart De Schutter · 2010
Earlier work this paper cites.
Multi-agent learning with policy prediction
Chongjie Zhang and Victor R. Lesser · 2010
Earlier work this paper cites.
The world of independent learners is not markovian
Guillaume J Laurent, Laëtitia Matignon, Le Fort-Piat, et al · 2011
Earlier work this paper cites.
An overview of recent progress in the study of distributed multi-agent coordination
Yongcan Cao, Wenwu Yu, Wei Ren, and Guanrong Chen · 2012
Earlier work this paper cites.
Necessity and sufficiency for the existence of a pure-strategy nash equilibrium
Jun-ichi Takeshita and Hidefumi Kawasaki · 2012
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Nash equilibria noncooperative games., sept 2013
Petr Šebek · 2013
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael I. Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning
Jakob N. Foerster, Yannis M. Assael, Nando de Freitas, and Shimon Whiteson · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemyslaw Debiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Later among the works it cites.
α \alpha -rank: Multi-agent evaluation by evolution
Shayegan Omidshafiei, Christos Papadimitriou, Georgios Piliouras, Karl Tuyls, Mark Rowland, Jean-Baptiste Lespiau, Wojciech M Czarnecki, Marc Lanctot, Julien Perolat, and Remi Munos · 2019
Later among the works it cites.
Open-ended learning in symmetric zero-sum games
David Balduzzi, Marta Garnelo, Yoram Bachrach, Wojciech Czarnecki, Julien Pérolat, Max Jaderberg, and Thore Graepel · 2019
Later among the works it cites.
Optimistic mirror descent in saddle-point problems: Going the extra (gradient) mile
Panayotis Mertikopoulos, Bruno Lecouat, Houssam Zenati, Chuan-Sheng Foo, Vijay Chandrasekhar, and Georgios Piliouras · 2019
Later among the works it cites.
Stable opponent shaping in differentiable games
Alistair Letcher, Jakob N. Foerster, David Balduzzi, Tim Rocktäschel, and Shimon Whiteson · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Learning multiagent communication with backpropagation
Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus · 2016
Cited alongside, same era.
Peng Peng, Ying Wen, Yaodong Yang, Quan Yuan, Zhenkun Tang, Haitao Long, and Jun Wang · 2017
Cited alongside, same era.
A survey of learning in multiagent environments: Dealing with non-stationarity
Pablo Hernandez-Leal, Michael Kaisers, Tim Baarslag, and Enrique Munoz de Cote · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinícius Flores Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Pérolat, David Silver, and Thore Graepel · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Later among the works it cites.
Policy optimization provably converges to nash equilibria in zero-sum linear quadratic games
Kaiqing Zhang, Zhuoran Yang, and Tamer Basar · 2019
Later among the works it cites.
Computing approximate equilibria in sequential adversarial games by exploitability descent
Edward Lockhart, Marc Lanctot, Julien Pérolat, Jean-Baptiste Lespiau, Dustin Morrill, Finbarr Timbers, and Karl Tuyls · 2019
Later among the works it cites.
On finding local nash equilibria (and only local nash equilibria) in zero-sum games
Eric V Mazumdar, Michael I Jordan, and S Shankar Sastry · 2019
Later among the works it cites.
Modelling bounded rationality in multi-agent interactions by generalized recursive reasoning
Ying Wen, Yaodong Yang, Rui Luo, and Jun Wang · 2019
Later among the works it cites.
QTRAN: learning to factorize with transformation for cooperative multi-agent reinforcement learning
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Hostallero, and Yung Yi · 2019
Later among the works it cites.
A collection of multi agent environments based on OpenAI gym, 2019
Anurag Koul · 2019
Later among the works it cites.
An overview of multi-agent reinforcement learning from game theoretical perspective
Yaodong Yang and Jun Wang · 2020
Later among the works it cites.
Smarts: Scalable multi-agent reinforcement learning training school for autonomous driving
Ming Zhou, Jun Luo, Julian Villela, Yaodong Yang, David Rusu, Jiayu Miao, Weinan Zhang, Montgomery Alban, Iman Fadakar, Zheng Chen, et al · 2020
Later among the works it cites.
On gradient-based learning in continuous games
Eric Mazumdar, Lillian J Ratliff, and S Shankar Sastry · 2020
Later among the works it cites.
A generalized training approach for multiagent learning
Paul Muller, Shayegan Omidshafiei, Mark Rowland, Karl Tuyls, Julien Pérolat, Siqi Liu, Daniel Hennes, Luke Marris, Marc Lanctot, Edward Hughes, Zhe Wang, Guy Lever, Nicolas Heess, Thore Graepel, and Rémi Munos · 2020
Later among the works it cites.
α \alpha α \alpha -rank: Practically scaling α \alpha -rank through stochastic optimisation
Yaodong Yang, Rasul Tutunov, Phu Sakulwongtana, and Haitham Bou Ammar · 2020
Later among the works it cites.
Multi-agent determinantal q-learning
Yaodong Yang, Ying Wen, Jun Wang, Liheng Chen, Kun Shao, David Mguni, and Weinan Zhang · 2020
Later among the works it cites.
Deep multi-agent reinforcement learning for decentralized continuous cooperative control, 2020
Christian Schroeder de Witt, Bei Peng, Pierre-Alexandre Kamienny, Philip Torr, Wendelin Böhmer, and Shimon Whiteson · 2020
Later among the works it cites.
Multiplayer support for the arcade learning environment
Justin K Terry and Benjamin Black · 2020
Later among the works it cites.
On the impossibility of global convergence in multi-loss optimization
Alistair Letcher · 2020
Later among the works it cites.
Bounds and dynamics for empirical game theoretic analysis
Karl Tuyls, Julien Perolat, Marc Lanctot, Edward Hughes, Richard Everett, Joel Z Leibo, Csaba Szepesvári, and Thore Graepel · 2020
Later among the works it cites.
Independent policy gradient methods for competitive reinforcement learning
Constantinos Daskalakis, Dylan J Foster, and Noah Golowich · 2020
Later among the works it cites.
Finite-time last-iterate convergence for multi-agent learning in games
Tianyi Lin, Zhengyuan Zhou, Panayotis Mertikopoulos, and Michael I. Jordan · 2020
Later among the works it cites.
Multi-agent trust region policy optimization
Hepeng Li and Haibo He · 2020
Later among the works it cites.
Dealing with non-stationarity in multi-agent reinforcement learning via trust region decomposition
Wenhao Li, Xiangfeng Wang, Bo Jin, Junjie Sheng, and Hongyuan Zha · 2021
Closest in time.