Fetching the paper…
Reading the bibliography…
Many real-world reinforcement learning tasks require multiple agents to make sequential decisions under the agents' interaction, where well-coordinated actions among the agents are crucial to achieve the target goal better at these tasks.
Auto-association by multilayer perceptrons and singular value decomposition
Hervé Bourlard and Yves Kamp · 1988
Earlier work this paper cites.
Stability properties of constrained queueing systems and scheduling policies for maximum throughput in multihop radio networks
Leandros Tassiulas and Anthony Ephremides · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
Autoencoders, minimum description length and helmholtz free energy
Geoffrey E Hinton and Richard S Zemel · 1994
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Multiagent systems: A survey from a machine learning perspective
Peter Stone and Manuela Veloso · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Wireless Communications: Principles and Practice
Theodore Rappaport · 2001
Earlier work this paper cites.
Coordinated reinforcement learning
Carlos Guestrin, Michail Lagoudakis, and R Parr · 2002
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2003
Earlier work this paper cites.
Contiki-a lightweight and flexible operating system for tiny networked sensors
Adam Dunkels, Bjorn Gronvall, and Thiemo Voigt · 2004
Earlier work this paper cites.
Computer networking: A top-down approach featuring the internet
James F Kurose · 2005
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Lucian Busoniu, Robert Babuska, and Bart De Schutter · 2008
Cited alongside, same era.
Complexity in wireless scheduling: Impact and tradeoffs
Yung Yi, Alexandre Proutière, and Mung Chiang · 2008
Cited alongside, same era.
A distributed csma algorithm for throughput and utility maximization in wireless networks
Libin Jiang and Jean Walrand · 2010
Cited alongside, same era.
Coordinating multi-agent reinforcement learning with limited communication
Chongjie Zhang and Victor Lesser · 2013
Cited alongside, same era.
Distributed learning for utility maximization over csma-based wireless multihop networks
Hyeryung Jang, Se-Young Yun, Jinwoo Shin, and Yung Yi · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cooperative multi-agent control using deep reinforcement learning
Jayesh K Gupta, Maxim Egorov, and Mykel Kochenderfer · 2017
Later among the works it cites.
Emergence of language with multi-agent games: learning to communicate with sequences of symbols
Serhii Havrylov and Ivan Titov · 2017
Later among the works it cites.
Multi-agent reinforcement learning in sequential social dilemmas
Joel Z Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, and Thore Graepel · 2017
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Later among the works it cites.
Neural adaptive video streaming with pensieve
Hongzi Mao, Ravi Netravali, and Mohammad Alizadeh · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning
Jakob Foerster, Ioannis Alexandros Assael, Nando de Freitas, and Shimon Whiteson · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
A concise introduction to decentralized POMDPs , volume 1
Frans A Oliehoek, Christopher Amato, et al · 2016
Cited alongside, same era.
Learning multiagent communication with backpropagation
Sainbayar Sukhbaatar, Rob Fergus, et al · 2016
Cited alongside, same era.
Igor Mordatch and Pieter Abbeel · 2017
Later among the works it cites.
Deep decentralized multi-task multi-agent RL under partial observability
Shayegan Omidshafiei, Jason Pazis, Christopher Amato, Jonathan P How, and John Vian · 2017
Later among the works it cites.
Lenient multi-agent deep reinforcement learning
Gregory Palmer, Karl Tuyls, Daan Bloembergen, and Rahul Savani · 2017
Later among the works it cites.
Multiagent bidirectionally-coordinated nets for learning to play starcraft combat games
Peng Peng, Quan Yuan, Ying Wen, Yaodong Yang, Zhenkun Tang, Haitao Long, and Jun Wang · 2017
Later among the works it cites.
Multiagent cooperation and competition with deep reinforcement learning
Ardi Tampuu, Tambet Matiisen, Dorian Kodelja, Ilya Kuzovkin, Kristjan Korjus, Juhan Aru, Jaan Aru, and Raul Vicente · 2017
Later among the works it cites.
Learning attentional communication for multi-agent cooperation
Jiechuan Jiang and Zongqing Lu · 2018
Later among the works it cites.
QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Later among the works it cites.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinícius Flores Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z. Leibo, Karl Tuyls, and Thore Graepel · 2018
Later among the works it cites.