Fetching the paper…
Reading the bibliography…
The main challenge of multiagent reinforcement learning is the difficulty of learning useful policies in the presence of other simultaneously learning agents whose changing behaviors jointly affect the environment's transition and reward dynamics.
Equilibrium points in n-person games
John F. Nash · 1950
Earlier work this paper cites.
Correlated equilibrium as an expression of bayesian rationality
Robert J. Aumann · 1987
Earlier work this paper cites.
Stochastic evolutionary game dynamics
Dean Foster and Peyton Young · 1990
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L. Littman · 1994
Earlier work this paper cites.
Friend-or-foe q-learning in general-sum games
Michael L. Littman · 2001
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Reinforcement learning to play an optimal nash equilibrium in team markov games
Xiaofeng Wang and Tuomas Sandholm · 2002
Earlier work this paper cites.
Correlated-Q learning
Amy Greenwald and Keith Hall · 2003
Earlier work this paper cites.
An algorithm for computing stochastically stable distributions with applications to multiagent learning in repeated games
John R. Wicks and Amy Greenwald · 2005
Earlier work this paper cites.
Convergence and no-regret in multiagent learning
Michael Bowling · 2005
Earlier work this paper cites.
Cyclic equilibria in markov games
Martin Zinkevich, Amy Greenwald, and Michael Littman · 2006
Earlier work this paper cites.
Multi-agent Reinforcement Learning: An Overview
Lucian Buşoniu, Robert Babuška, and Bart De Schutter · 2010
Earlier work this paper cites.
Multi-agent learning with policy prediction
Chongjie Zhang and Victor R. Lesser · 2010
Earlier work this paper cites.
Game Theory and Multi-agent Reinforcement Learning
Ann Nowé, Peter Vrancx, and Yann-Michaël De Hauwere · 2012
Earlier work this paper cites.
Random Perturbations of Dynamical Systems
M.I. Freidlin, J. Szücs, and A.D. Wentzell · 2012
Earlier work this paper cites.
Distributed dynamic reinforcement of efficient outcomes in multiagent coordination and network formation
Georgios Chasparis and Jeff S. Shamma · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Earlier work this paper cites.
Opponent modeling in deep reinforcement learning
He He, Jordan Boyd-Graber, Kevin Kwok, and Hal Daumé III · 2016
Earlier work this paper cites.
Opponent modeling in deep reinforcement learning
He He, Jordan L. Boyd-Graber, Kevin Kwok, and Hal Daumé III · 2016
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Variational inference: A review for statisticians
David M. Blei, Alp Kucukelbir, and Jon D. McAuliffe · 2017
Cited alongside, same era.
A survey of learning in multiagent environments: Dealing with non-stationarity
Pablo Hernandez-Leal, Michael Kaisers, Tim Baarslag, and Enrique Munoz de Cote · 2017
Cited alongside, same era.
Multi-agent reinforcement learning in sequential social dilemmas
Joel Z. Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, and Thore Graepel · 2017
Cited alongside, same era.
Meta-learning with implicit gradients
Aravind Rajeswaran, Chelsea Finn, Sham M Kakade, and Sergey Levine · 2019
Later among the works it cites.
Soft actor-critic for discrete action settings
Petros Christodoulou · 2019
Later among the works it cites.
Learning to teach in cooperative multiagent reinforcement learning
Shayegan Omidshafiei, Dong-Ki Kim, Miao Liu, Gerald Tesauro, Matthew Riemer, Christopher Amato, Murray Campbell, and Jonathan P. How · 2019
Later among the works it cites.
Policy distillation and value matching in multiagent reinforcement learning
Samir Wadhwania, Dong-Ki Kim, Shayegan Omidshafiei, and Jonathan P. How · 2019
Later among the works it cites.
Probabilistic recursive reasoning for multi-agent reinforcement learning
Ying Wen, Yaodong Yang, Rui Luo, Jun Wang, and Wei Pan · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning with opponent-learning awareness
Jakob Foerster, Richard Y. Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch · 2018
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Cited alongside, same era.
Mean field multi-agent reinforcement learning
Yaodong Yang, Rui Luo, Minne Li, Ming Zhou, Weinan Zhang, and Jun Wang · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Maruan Al-Shedivat, Trapit Bansal, Yura Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Magent: A many-agent reinforcement learning platform for artificial collective intelligence
Lianmin Zheng, Jiacheng Yang, Han Cai, Ming Zhou, Weinan Zhang, Jun Wang, and Yong Yu · 2018
Cited alongside, same era.
Social influence as intrinsic motivation for multi-agent deep reinforcement learning
Natasha Jaques, Angeliki Lazaridou, Edward Hughes, Caglar Gulcehre, Pedro Ortega, Dj Strouse, Joel Z. Leibo, and Nando De Freitas · 2019
Later among the works it cites.
Evolving intrinsic motivations for altruistic behavior
Jane X. Wang, Edward Hughes, Chrisantha Fernando, Wojciech M. Czarnecki, Edgar A. Duéñez Guzmán, and Joel Z. Leibo · 2019
Later among the works it cites.
Achieving cooperation through deep multiagent reinforcement learning in sequential prisoner’s dilemmas
Weixun Wang, Jianye Hao, Yixi Wang, and Matthew Taylor · 2019
Later among the works it cites.
Learning latent representations to influence multi-agent interaction
Annie Xie, Dylan Losey, Ryan Tolsma, Chelsea Finn, and Dorsa Sadigh · 2020
Later among the works it cites.
Learning hierarchical teaching policies for cooperative agents
Dong-Ki Kim, Miao Liu, Shayegan Omidshafiei, Sebastian Lopez-Cot, Matthew Riemer, Golnaz Habibi, Gerald Tesauro, Sami Mourad, Murray Campbell, and Jonathan P. How · 2020
Later among the works it cites.
Learning to incentivize other learning agents
Jiachen Yang, Ang Li, Mehrdad Farajtabar, Peter Sunehag, Edward Hughes, and Hongyuan Zha · 2020
Later among the works it cites.
Varibad: A very good method for bayes-adaptive deep rl via meta-learning
Luisa Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze, Yarin Gal, Katja Hofmann, and Shimon Whiteson · 2020
Later among the works it cites.
A policy gradient algorithm for learning to learn in multiagent reinforcement learning
Dong Ki Kim, Miao Liu, Matthew D Riemer, Chuangchuang Sun, Marwa Abdulhai, Golnaz Habibi, Sebastian Lopez-Cot, Gerald Tesauro, and Jonathan How · 2021
Later among the works it cites.
Influencing towards stable multi-agent interactions
Woodrow Zhouyuan Wang, Andy Shih, Annie Xie, and Dorsa Sadigh · 2021
Later among the works it cites.
Model-free opponent shaping
Christopher Lu, Timon Willi, Christian A Schroeder De Witt, and Jakob Foerster · 2022
Closest in time.
The complexity of markov equilibrium in stochastic games
Constantinos Daskalakis, Noah Golowich, and Kaiqing Zhang · 2022
Closest in time.
Continuous-time meta-learning with forward mode differentiation
Tristan Deleu, David Kanaa, Leo Feng, Giancarlo Kerg, Yoshua Bengio, Guillaume Lajoie, and Pierre-Luc Bacon · 2022
Closest in time.
The good shepherd: An oracle agent for mechanism design
Jan Balaguer, Raphael Koster, Christopher Summerfield, and Andrea Tacchetti · 2022
Closest in time.
Adaptive incentive design with multi-agent meta-gradient reinforcement learning
Jiachen Yang, Ethan Wang, Rakshit Trivedi, Tuo Zhao, and Hongyuan Zha · 2022
Closest in time.