Fetching the paper…
Reading the bibliography…
Non-stationarity is one thorny issue in cooperative multi-agent reinforcement learning (MARL).
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Temporal difference learning and TD-Gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
Sergiu Hart and Andreu Mas-Colell · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
Learning fair policies in decentralized cooperative multi-agent reinforcement learning
Matthieu Zimmer, Claire Glanois, Umer Siddique, and Paul Weng · 2004
Earlier work this paper cites.
Dynamic programming for partially observable stochastic games
Eric A Hansen, Daniel S Bernstein, and Shlomo Zilberstein · 2004
Earlier work this paper cites.
Revisiting parameter sharing in multi-agent deep reinforcement learning
Justin K Terry, Nathaniel Grammel, Ananth Hari, Luis Santos, and Benjamin Black · 2005
Earlier work this paper cites.
Optimal and approximate Q-value functions for decentralized POMDPs
Frans A Oliehoek, Matthijs TJ Spaan, and Nikos Vlassis · 2008
Earlier work this paper cites.
PettingZoo: Gym for multi-agent reinforcement learning
Justin K Terry, Benjamin Black, Mario Jayakumar, Ananth Hari, Luis Santos, Clemens Dieffendahl, Niall L Williams, Yashas Lokesh, Ryan Sullivan, Caroline Horsch, and Praveen Ravi · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Online learning with switching costs and other adaptive adversaries
Nicolo Cesa-Bianchi, Ofer Dekel, and Ohad Shamir · 2013
Earlier work this paper cites.
Stochastic multi-armed-bandit problem with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Chapter 4. Approximation to rayes risk in repeated play
James Hannan · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Earlier work this paper cites.
Cooperative multi-agent control using deep reinforcement learning
Jayesh K Gupta, Maxim Egorov, and Mykel Kochenderfer · 2017
Earlier work this paper cites.
A survey of learning in multiagent environments: Dealing with non-stationarity
Pablo Hernandez-Leal, Michael Kaisers, Tim Baarslag, and Enrique Munoz de Cote · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Stabilising experience replay for deep multi-agent reinforcement learning
Jakob N. Foerster, Nantas Nardelli, Gregory Farquhar, Triantafyllos Afouras, Philip H. S. Torr, Pushmeet Kohli, and Shimon Whiteson · 2017
Earlier work this paper cites.
Cooperative multi-agent control using deep reinforcement learning
Jayesh K Gupta, Maxim Egorov, and Mykel Kochenderfer · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Yuhuai Wu, Elman Mansimov, Roger B Grosse, Shun Liao, and Jimmy Ba · 2017
Earlier work this paper cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Maruan Al-Shedivat, Trapit Bansal, Yura Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Earlier work this paper cites.
Machine theory of mind
Neil Rabinowitz, Frank Perbet, Francis Song, Chiyuan Zhang, SM Ali Eslami, and Matthew Botvinick · 2018
Cited alongside, same era.
Modeling others using oneself in multi-agent reinforcement learning
Roberta Raileanu, Emily Denton, Arthur Szlam, and Rob Fergus · 2018
Cited alongside, same era.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Maruan Al-Shedivat, Trapit Bansal, Yura Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2018
Cited alongside, same era.
Pratik Gajane, Ronald Ortner, and Peter Auer · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Multi-agent trust region policy optimization
Hepeng Li and Haibo He · 2020
Later among the works it cites.
Variational regret bounds for reinforcement learning
Ronald Ortner, Pratik Gajane, and Peter Auer · 2020
Later among the works it cites.
Intention propagation for multi-agent reinforcement learning
Chao Qu, Hui Li, Chang Liu, Junwu Xiong, James Zhang, Wei Chu, Yuan Qi, and Le Song · 2020
Later among the works it cites.
Revisiting parameter sharing in multi-agent deep reinforcement learning
Justin K Terry, Nathaniel Grammel, Ananth Hari, Luis Santos, and Benjamin Black · 2020
Later among the works it cites.
Mirror descent policy optimization
Manan Tomar, Lior Shani, Yonathan Efroni, and Mohammad Ghavamzadeh · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Trust-pcl: An off-policy trust region method for continuous control
Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans · 2018
Cited alongside, same era.
Machine theory of mind
Neil Rabinowitz, Frank Perbet, Francis Song, Chiyuan Zhang, SM Ali Eslami, and Matthew Botvinick · 2018
Cited alongside, same era.
Modeling others using oneself in multi-agent reinforcement learning
Roberta Raileanu, Emily Denton, Arthur Szlam, and Rob Fergus · 2018
Cited alongside, same era.
Provably efficient q-learning with low switching cost
Yu Bai, Tengyang Xie, Nan Jiang, and Yuxiang Wang · 2019
Cited alongside, same era.
Emergent tool use from multi-agent autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch · 2019
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, G. Brockman, Brooke Chan, Vicki Cheung, Przemyslaw Debiak, Christy Dennison, D. Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, R. Józefowicz, Scott Gray, C. Olsson, Jakub W. Pachocki, M. Petrov, Henrique Pondé de Oliveira Pinto, Jonathan Raiman, Tim Salimans, Jeremy Schlatter, J. Schneider, S. Sidor, Ilya Sutskever, Jie Tang, F. Wolski, and Susan Zhang · 2019
Cited alongside, same era.
Non-stationary reinforcement learning: The blessing of (more) optimism
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2019
Cited alongside, same era.
Later among the works it cites.
Towards playing full MOBA games with deep reinforcement learning
Deheng Ye, Guibin Chen, W. Zhang, Sheng Chen, Bo Yuan, B. Liu, J. Chen, Z. Liu, Fuhao Qiu, Hongsheng Yu, Yinyuting Yin, Bei Shi, L. Wang, Tengfei Shi, Qiang Fu, Wei Yang, L. Huang, and Wei Liu · 2020
Later among the works it cites.
MOPO: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Is independent learning all you need in the starcraft multi-agent challenge?
Christian Schroeder de Witt, Tarun Gupta, Denys Makoviichuk, Viktor Makoviychuk, Philip HS Torr, Mingfei Sun, and Shimon Whiteson · 2020
Later among the works it cites.
Multi-agent trust region policy optimization
Hepeng Li and Haibo He · 2020
Later among the works it cites.
Variational regret bounds for reinforcement learning
Ronald Ortner, Pratik Gajane, and Peter Auer · 2020
Later among the works it cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Lior Shani, Yonathan Efroni, and Shie Mannor · 2020
Later among the works it cites.
Arena: A general evaluation platform and building toolkit for multi-agent intelligence
Yuhang Song, Jianyi Wang, Thomas Lukasiewicz, Zhenghua Xu, Mai Xu, Zihan Ding, and Lianlong Wu · 2020
Later among the works it cites.
Mirror descent policy optimization
Manan Tomar, Lior Shani, Yonathan Efroni, and Mohammad Ghavamzadeh · 2020
Later among the works it cites.
Learning latent representations to influence multi-agent interaction
Annie Xie, Dylan Losey, Ryan Tolsma, Chelsea Finn, and Dorsa Sadigh · 2020
Later among the works it cites.
A provably efficient algorithm for linear markov decision process with low switching cost
Minbo Gao, Tianle Xie, Simon S Du, and Lin F Yang · 2021
Closest in time.
Adaptive learning in continuous games: Optimal regret bounds and convergence to nash equilibrium
Yu-Guan Hsieh, Kimon Antonakopoulos, and Panayotis Mertikopoulos · 2021
Closest in time.
Deep implicit coordination graphs for multi-agent reinforcement learning
Sheng Li, Jayesh K Gupta, Peter Morales, Ross Allen, and Mykel J Kochenderfer · 2021
Closest in time.
Near-optimal model-free reinforcement learning in non-stationary episodic mdps
Weichao Mao, Kaiqing Zhang, Ruihao Zhu, David Simchi-Levi, and Tamer Başar · 2021
Closest in time.
A survey of reinforcement learning algorithms for dynamically varying environments
Sindhu Padakandla · 2021
Closest in time.
The surprising effectiveness of MAPPO in cooperative, multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Yu Wang, Alexandre M. Bayen, and Yi Wu · 2021
Closest in time.
A kernel-based approach to non-stationary reinforcement learning in metric spaces
Omar Darwiche Domingues, Pierre Ménard, Matteo Pirotta, Emilie Kaufmann, and Michal Valko · 2021
Closest in time.
Noisy-MAPPO: Noisy advantage values for cooperative multi-agent actor-critic methods
Siyue Hu and Jian Hu · 2021
Closest in time.
Adaptive learning rates for multi-agent reinforcement learning, 2021
Jiechuan Jiang and Zongqing Lu · 2021
Closest in time.
A policy gradient algorithm for learning to learn in multiagent reinforcement learning
Dong Ki Kim, Miao Liu, Matthew D Riemer, Chuangchuang Sun, Marwa Abdulhai, Golnaz Habibi, Sebastian Lopez-Cot, Gerald Tesauro, and Jonathan How · 2021
Closest in time.
Trust region policy optimisation in multi-agent reinforcement learning
Jakub Grudzien Kuba, Ruiqing Chen, Munning Wen, Ying Wen, Fanglei Sun, Jun Wang, and Yaodong Yang · 2021
Closest in time.
Near-optimal model-free reinforcement learning in non-stationary episodic mdps
Weichao Mao, Kaiqing Zhang, Ruihao Zhu, David Simchi-Levi, and Tamer Başar · 2021
Closest in time.
A game-theoretic approach to multi-agent trust region optimization
Ying Wen, Hui Chen, Yaodong Yang, Zheng Tian, Minne Li, Xu Chen, and Jun Wang · 2021
Closest in time.
The surprising effectiveness of MAPPO in cooperative, multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Yu Wang, Alexandre M. Bayen, and Yi Wu · 2021
Closest in time.
Learning to schedule multi-numa virtual machines via reinforcement learning
Junjie Sheng, Yiqiu Hu, Wenli Zhou, Lei Zhu, Bo Jin, Jun Wang, and Xiangfeng Wang · 2022
Closest in time.