Fetching the paper…
Reading the bibliography…
Learning to communicate in order to share state information is an active problem in the area of multi-agent reinforcement learning (MARL).
On-line Q-learning using connectionist systems . Vol. 37
Gavin A Rummery and Mahesan Niranjan. 1994 · 1994
Earlier work this paper cites.
Introduction to Reinforcement Learning (1st ed.)
Richard S. Sutton and Andrew G. Barto. 1998 · 1998
Earlier work this paper cites.
All learning is local: Multi-agent learning in global reward games
Yu-Han Chang, Tracey Ho, and Leslie P Kaelbling. 2004 · 2004
Earlier work this paper cites.
A Comprehensive Survey of Multiagent Reinforcement Learning
L. Busoniu, R. Babuska, and B. De Schutter. 2008 · 2008
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013 · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms. In International conference on machine learning . PMLR, 387–395
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. 2014 · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Deep Reinforcement Learning with Double Q-learning
Hado van Hasselt, Arthur Guez, and David Silver. 2015 · 2015
Earlier work this paper cites.
Dueling Network Architectures for Deep Reinforcement Learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, and Nando de Freitas. 2015 · 2015
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning. In Advances in neural information processing systems . 2137–2145
Jakob Foerster, Ioannis Alexandros Assael, Nando De Freitas, and Shimon Whiteson. 2016 · 2016
Earlier work this paper cites.
Learning to play guess who? and inventing a grounded language as a consequence
Emilio Jorge, Mikael Kågebäck, Fredrik D Johansson, and Emil Gustavsson. 2016 · 2016
Cited alongside, same era.
Asynchronous Methods for Deep Reinforcement Learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016 · 2016
Cited alongside, same era.
A concise introduction to decentralized POMDPs . Vol. 1
Frans A Oliehoek, Christopher Amato, et al · 2016
Cited alongside, same era.
Learning Multiagent Communication with Backpropagation. In Advances in Neural Information Processing Systems , D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett (Eds.), Vol. 29. Curran Associates, Inc
Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus. 2016 · 2016
Cited alongside, same era.
Counterfactual multi-agent policy gradients. In Thirty-second AAAI conference on artificial intelligence
Jakob N Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson. 2018 · 2018
Later among the works it cites.
Learning Attentional Communication for Multi-Agent Cooperation
Jiechuan Jiang and Zongqing Lu. 2018 · 2018
Later among the works it cites.
RLlib: Abstractions for Distributed Reinforcement Learning. In International Conference on Machine Learning (ICML)
Eric Liang, Richard Liaw, Robert Nishihara, Philipp Moritz, Roy Fox, Ken Goldberg, Joseph E. Gonzalez, Michael I. Jordan, and Ion Stoica. 2018 · 2018
Later among the works it cites.
Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning. In International Conference on Machine Learning . 3040–3049
Natasha Jaques, Angeliki Lazaridou, Edward Hughes, Caglar Gulcehre, Pedro Ortega, Dj Strouse, Joel Z Leibo, and Nando De Freitas. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jakob Foerster, Nantas Nardelli, Gregory Farquhar, Triantafyllos Afouras, Philip HS Torr, Pushmeet Kohli, and Shimon Whiteson. 2017 · 2017
Cited alongside, same era.
Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Cited alongside, same era.
Hangyu Mao, Zhibo Gong, Yan Ni, and Zhen Xiao. 2017 · 2017
Cited alongside, same era.
Emergence of Grounded Compositional Language in Multi-Agent Populations
Igor Mordatch and Pieter Abbeel. 2017 · 2017
Cited alongside, same era.
Multiagent Bidirectionally-Coordinated Nets for Learning to Play StarCraft Combat Games
Peng Peng, Quan Yuan, Ying Wen, Yaodong Yang, Zhenkun Tang, Haitao Long, and Jun Wang. 2017 · 2017
Cited alongside, same era.
TarMAC: Targeted Multi-Agent Communication
Abhishek Das, Théophile Gervet, Joshua Romoff, Dhruv Batra, Devi Parikh, Michael Rabbat, and Joelle Pineau. 2018 · 2018
Cited alongside, same era.
Ryan Lowe, Jakob Foerster, Y-Lan Boureau, Joelle Pineau, and Yann Dauphin. 2019 · 2019
Later among the works it cites.
Hierarchical Multi-Agent Deep Reinforcement Learning to Develop Long-Term Coordination. In Proceedings of the 34th ACM/SIGAPP Symposium on Applied Computing (Limassol, Cyprus) (SAC ’19) . Association for Computing Machinery, New York, NY, USA, 922–929
Marie Ossenkopf, Mackenzie Jorgensen, and Kurt Geihs. 2019 · 2019
Later among the works it cites.
Learning individually inferred communication for multi-agent cooperation
Ziluo Ding, Tiejun Huang, and Zongqing Lu. 2020 · 2020
Closest in time.
Multi-agent actor centralized-critic with communication
David Simões, Nuno Lau, and Luís Paulo Reis. 2020 · 2020
Closest in time.
Learning to Communicate with Multi-agent Reinforcement Learning Using Value-Decomposition Networks. In Advances on P2P, Parallel, Grid, Cloud and Internet Computing , Leonard Barolli, Peter Hellinckx, and Juggapong Natwichai (Eds.). Springer International Publishing, Cham, 736–745
Simon Vanneste, Astrid Vanneste, Stig Bosmans, Siegfried Mercelis, and Peter Hellinckx. 2020 · 2020
Closest in time.
Value-Decomposition Networks For Cooperative Multi-Agent Learning Based On Team Reward. In Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems (Stockholm, Sweden) (AAMAS ’18) . International Foundation for Autonomous Agents and Multiagent Systems, Richland, SC, 2085–2087
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z. Leibo, Karl Tuyls, and Thore Graepel. 2018 · 2087
Closest in time.