Fetching the paper…
Reading the bibliography…
Many real-world tasks involve multiple agents with partial observability and limited communication.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, Long-Ji · 1992
Earlier work this paper cites.
Q-learning
Watkins, Christopher JCH and Dayan, Peter · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Tan, Ming · 1993
Earlier work this paper cites.
Stable function approximation in dynamic programming
Gordon, Geoffrey J · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Multitask learning
Caruana, Rich · 1998
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Claus, Caroline and Boutilier, Craig · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, Leslie Pack, Littman, Michael L, and Cassandra, Anthony R · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, Richard S and Barto, Andrew G · 1998
Earlier work this paper cites.
An algorithm for distributed reinforcement learning in cooperative multi-agent systems
Lauer, Martin and Riedmiller, Martin · 2000
Earlier work this paper cites.
Learning to cooperate via policy search
Peshkin, Leonid, Kim, Kee-Eung, Meuleau, Nicolas, and Kaelbling, Leslie Pack · 2000
Earlier work this paper cites.
Multi-agent systems by incremental gradient reinforcement learning
Dutech, Alain, Buffet, Olivier, and Charpillet, François · 2001
Earlier work this paper cites.
The complexity of decentralized control of markov decision processes
Bernstein, Daniel S, Givan, Robert, Immerman, Neil, and Zilberstein, Shlomo · 2002
Earlier work this paper cites.
Multiagent learning using a variable learning rate
Bowling, Michael and Veloso, Manuela · 2002
Earlier work this paper cites.
Reinforcement learning of coordination in cooperative multi-agent systems
Kapetanakis, Spiros and Kudenko, Daniel · 2002
Earlier work this paper cites.
Multitask reinforcement learning on the distribution of mdps
Tanaka, Fumihide and Yamamura, Masayuki · 2003
Earlier work this paper cites.
Probabilistic policy reuse in a reinforcement learning agent
Fernández, Fernando and Veloso, Manuela · 2006
Cited alongside, same era.
Predicting and preventing coordination problems in cooperative Q-learning systems
Fulda, Nancy and Ventura, Dan · 2007
Cited alongside, same era.
Hysteretic Q-learning: an algorithm for decentralized reinforcement learning in cooperative multi-agent teams
Matignon, Laëtitia, Laurent, Guillaume J, and Le Fort-Piat, Nadine · 2007
Cited alongside, same era.
Solving deep memory POMDPs with recurrent policy gradients
Wierstra, Daan, Foerster, Alexander, Peters, Jan, and Schmidhuber, Juergen · 2007
Cited alongside, same era.
Multi-task reinforcement learning: a hierarchical bayesian approach
Wilson, Aaron, Fern, Alan, Ray, Soumya, and Tadepalli, Prasad · 2007
Cited alongside, same era.
Optimal and approximate q-value functions for decentralized POMDPs
Rollout sampling policy iteration for decentralized POMDPs
Wu, Feng, Zilberstein, Shlomo, and Chen, Xiaoping · 2012
Later among the works it cites.
Multiagent-based reinforcement learning for optimal reactive power dispatch
Xu, Yinliang, Zhang, Wei, Liu, Wenxin, and Ferrese, Frank · 2012
Later among the works it cites.
Sample complexity of multi-task reinforcement learning
Brunskill, Emma and Li, Lihong · 2013
Later among the works it cites.
Transfer learning in multi-agent systems through parallel transfer
Taylor, Adam, Dusparic, Ivana, Galván-López, Edgar, Clarke, Siobhán, and Cahill, Vinny · 2013
Later among the works it cites.
Monte-carlo expectation maximization for decentralized POMDPs
Wu, Feng, Zilberstein, Shlomo, and Jennings, Nicholas R · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Oliehoek, Frans A, Spaan, Matthijs TJ, Vlassis, Nikos A, et al · 2008
Cited alongside, same era.
Learning and transferring roles in multi-agent reinforcement
Wilson, Aaron, Fern, Alan, Ray, Soumya, and Tadepalli, Prasad · 2008
Cited alongside, same era.
Incremental policy generation for finite-horizon DEC-POMDPs
Amato, Christopher, Dibangoye, Jilles Steeve, and Zilberstein, Shlomo · 2009
Cited alongside, same era.
Transfer learning for reinforcement learning domains: A survey
Taylor, Matthew E and Stone, Peter · 2009
Cited alongside, same era.
Transfer learning
Torrey, Lisa and Shavlik, Jude · 2009
Cited alongside, same era.
Multi-agent reinforcement learning: An overview
Buşoniu, Lucian, Babuška, Robert, and De Schutter, Bart · 2010
Cited alongside, same era.
A survey on transfer learning
Pan, Sinno Jialin and Yang, Qiang · 2010
Cited alongside, same era.
A robust approach for multi-agent natural resource allocation based on stochastic optimization algorithms
Barbalios, Nikos and Tzionas, Panagiotis · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Later among the works it cites.
Deep recurrent Q-learning for partially observable MDPs
Hausknecht, Matthew and Stone, Peter · 2015
Later among the works it cites.
Distilling the knowledge in a neural network
Hinton, Geoffrey, Vinyals, Oriol, and Dean, Jeff · 2015
Later among the works it cites.
Stick-breaking policy learning in Dec-POMDPs
Liu, Miao, Amato, Christopher, Liao, Xuejun, Carin, Lawrence, and How, Jonathan P · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Later among the works it cites.
Rusu, Andrei A, Colmenarejo, Sergio Gomez, Gulcehre, Caglar, Desjardins, Guillaume, Kirkpatrick, James, Pascanu, Razvan, Mnih, Volodymyr, Kavukcuoglu, Koray, and Hadsell, Raia · 2015
Later among the works it cites.
Learning to communicate to solve riddles with deep distributed recurrent Q-networks
Foerster, Jakob N, Assael, Yannis M, de Freitas, Nando, and Whiteson, Shimon · 2016
Later among the works it cites.
Learning for decentralized control of multiagent systems in large partially observable stochastic environments
Liu, Miao, Amato, Christopher, Anesta, Emily, Griffith, J. Daniel, and How, Jonathan P · 2016
Later among the works it cites.
A Concise Introduction to Decentralized POMDPs
Oliehoek, Frans A. and Amato, Christopher · 2016
Later among the works it cites.
Stabilising experience replay for deep multi-agent reinforcement learning
Foerster, Jakob, Nardelli, Nantas, Farquhar, Gregory, Torr, Philip, Kohli, Pushmeet, Whiteson, Shimon, et al · 2017
Closest in time.