Fetching the paper…
Reading the bibliography…
Many real world tasks require multiple agents to work together.
Multi-agent reinforcement learning: Independent vs. cooperative agents
M. Tan · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
M. L. Littman · 1994
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Policy Gradient Methods for Reinforcement Learning with Function Approximation
R. S. Sutton, D. Mcallester, S. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
Dynamic Programming For Partially Observable Stochastic Games
E. A. Hansen, D. S. Bernstein, and S. Zilberstein · 2004
Earlier work this paper cites.
A multi-agent approach to cooperative traffic management and route guidance
J. L. Adler, G. Satapathy, V. Manikonda, B. Bowles, and V. J. Blue · 2005
Earlier work this paper cites.
Multi-Agent Shared Hierarchy Reinforcement Learning
N. Mehta, P. Tadepalli, and C. Science · 2005
Earlier work this paper cites.
Decentralized reinforcement learning control of a robotic manipulator
L. Buşoniu, B. De Schutter, and R. Babuška · 2006
Earlier work this paper cites.
Central pattern generators for locomotion control in animals and robots: A review
A. J. Ijspeert · 2008
Earlier work this paper cites.
Optimal and approximate Q-value functions for decentralized POMDPs
F. A. Oliehoek, M. T. Spaan, and N. Vlassis · 2008
Earlier work this paper cites.
Double Q-Learning
H. van Hasselt · 2010
Cited alongside, same era.
Optimal control in microgrid using multi-agent reinforcement learning
F. D. Li, M. Wu, Y. He, and X. Chen · 2012
Cited alongside, same era.
Deterministic Policy Gradient Algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. a. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Cited alongside, same era.
A multi-agent framework for packet routing in wireless sensor networks
Shape-based compliant control with variable coordination centralization on a snake robot
J. Whitman, F. Ruscelli, M. Travers, and H. Choset · 2016
Later among the works it cites.
Categorical Reparameterization with Gumbel-Softmax
E. Jang, S. Gu, and B. Poole · 2017
Later among the works it cites.
Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch · 2017
Later among the works it cites.
Multi-agent Double Deep Q-Networks
D. Simões, N. Lau, and L. P. Reis · 2017
Later among the works it cites.
Counterfactual Multi-Agent Policy Gradients
J. Foerster, G. Farquhar, T. Afouras, N. Nardelli, and S. Whiteson · 2018
Later among the works it cites.
Addressing Function Approximation Error in Actor-Critic Methods
S. Fujimoto, H. van Hoof, and D. Meger · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Ye, M. Zhang, and Y. Yang · 2015
Cited alongside, same era.
Openai gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Cited alongside, same era.
Deep Reinforcement Learning with Double Q-learning
H. van Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
Emergence of Grounded Compositional Language in Multi-Agent Populations
I. Mordatch and P. Abbeel · 2018
Later among the works it cites.
Distributed learning for the decentralized control of articulated mobile robots
G. Sartoretti, Y. Shi, W. Paivine, M. Travers, and H. Choset · 2018
Later among the works it cites.
Robust Multi-Agent Reinforcement Learning via Minimax Deep Deterministic Policy Gradient
S. Li, Y. Wu, F. Fang, and S. Russell · 2019
Closest in time.