Fetching the paper…
Reading the bibliography…
In multi-agent reinforcement learning, the inherent non-stationarity of the environment caused by other agents' actions posed significant difficulties for an agent to learn a good policy independently.
Stimulus and response generalization: A stochastic model relating generalization to distance in psychological space
Roger N Shepard · 1957
Earlier work this paper cites.
Exemplar-based model of social judgment
Eliot R Smith and Michael A Zarate · 1992
Earlier work this paper cites.
Similarity as transformation
Ulrike Hahn, Nick Chater, and Lucy B Richardson · 2003
Earlier work this paper cites.
Wasserstein barycenter and its application to texture mixing
Julien Rabin, Gabriel Peyré, Julie Delon, and Marc Bernot · 2011
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
Jakob Foerster, Ioannis Alexandros Assael, Nando De Freitas, and Shimon Whiteson · 2016
Earlier work this paper cites.
Opponent modeling in deep reinforcement learning
He He, Jordan Boyd-Graber, Kevin Kwok, and Hal Daumé III · 2016
Earlier work this paper cites.
Identifying and tracking switching, non-stationary opponents: A bayesian approach
Pablo Hernandez-Leal, Matthew E Taylor, Benjamin Rosman, L Enrique Sucar, and Enrique Munoz De Cote · 2016
Earlier work this paper cites.
Learning multiagent communication with backpropagation
Sainbayar Sukhbaatar, Rob Fergus, et al · 2016
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Maruan Al-Shedivat, Trapit Bansal, Yuri Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2018
Cited alongside, same era.
Autonomous agents modelling other agents: A comprehensive survey and open problems
Stefano V Albrecht and Peter Stone · 2018
Cited alongside, same era.
Generative modeling using the sliced wasserstein distance
Ishan Deshpande, Ziyu Zhang, and Alexander G. Schwing · 2018
Cited alongside, same era.
Learning against non-stationary agents with opponent modelling and deep reinforcement learning
Richard Everett and Stephen Roberts · 2018
Cited alongside, same era.
Learning with opponent-learning awareness
Modeling others using oneself in multi-agent reinforcement learning
Roberta Raileanu, Emily Denton, Arthur Szlam, and Rob Fergus · 2018
Later among the works it cites.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder De Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Later among the works it cites.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinícius Flores Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2018
Later among the works it cites.
A deep bayesian policy reuse approach against non-stationary agents
Yan Zheng, Zhaopeng Meng, Jianye Hao, Zongzhang Zhang, Tianpei Yang, and Changjie Fan · 2018
Later among the works it cites.
Learning actionable representations with goal conditioned policies
Dibya Ghosh, Abhishek Gupta, and Sergey Levine · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jakob Foerster, Richard Y Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch · 2018
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Cited alongside, same era.
Learning policy representations in multiagent systems
Aditya Grover, Maruan Al-Shedivat, Jayesh K. Gupta, Yuri Burda, and Harrison Edwards · 2018
Cited alongside, same era.
A deep policy inference q-network for multi-agent systems
Zhang-Wei Hong, Shih-Yang Su, Tzu-Yun Shann, Yi-Hsiang Chang, and Chun-Yi Lee · 2018
Cited alongside, same era.
Learning attentional communication for multi-agent cooperation
Jiechuan Jiang and Zongqing Lu · 2018
Cited alongside, same era.
Emergence of grounded compositional language in multi-agent populations
Igor Mordatch and Pieter Abbeel · 2018
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
A policy gradient algorithm for learning to learn in multiagent reinforcement learning
Dong-Ki Kim, Miao Liu, Matthew Riemer, Chuangchuang Sun, Marwa Abdulhai, Golnaz Habibi, Sebastian Lopez-Cot, Gerald Tesauro, and Jonathan P How · 2020
Later among the works it cites.
Google research football: A novel reinforcement learning environment
Karol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajac, Olivier Bachem, Lasse Espeholt, Carlos Riquelme, Damien Vincent, Marcin Michalski, Olivier Bousquet, et al · 2020
Later among the works it cites.
Qplex: Duplex dueling multi-agent q-learning
Jianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu, and Chongjie Zhang · 2020
Later among the works it cites.
Fop: Factorizing optimal joint policy of maximum-entropy multi-agent reinforcement learning
Tianhao Zhang, Yueheng Li, Chen Wang, Guangming Xie, and Zongqing Lu · 2021
Closest in time.