Fetching the paper…
Reading the bibliography…
In this paper, we explore using deep reinforcement learning for problems with multiple agents.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine learning
1992
Earlier work this paper cites.
M. L. Littman, “Markov games as a framework for multi-agent reinforcement learning,” in Machine Learning Proceedings 1994
1994
Earlier work this paper cites.
MIT press Cambridge, 1998
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction · 1998
Earlier work this paper cites.
J. Hu and M. P. Wellman, “Nash q-learning for general-sum stochastic games,” Journal of machine learning research
2003
Earlier work this paper cites.
B. C. Da Silva, E. W. Basso, A. L. Bazzan, and P. M. Engel, “Dealing with non-stationary environments using context detection,” in Proceedings of the 23rd international conference on Machine learning
2006
Earlier work this paper cites.
L. Busoniu, R. Babuska, and B. De Schutter, “Multi-agent reinforcement learning: A survey,” in Control, Automation, Robotics and Vision, 2006. ICARCV’06. 9th International Conference on
2006
Earlier work this paper cites.
R. S. Sutton, A. Koop, and D. Silver, “On the role of tracking in stationary environments,” in Proceedings of the 24th international conference on Machine learning
2007
Earlier work this paper cites.
V. Conitzer and T. Sandholm, “Awesome: A general multiagent learning algorithm that converges in self-play and learns a best response against stationary opponents,” Machine Learning
2007
Earlier work this paper cites.
J. Enright and P. R. Wurman, “Optimization and coordinated autonomy in mobile fulfillment systems.,” 2011
2011
Earlier work this paper cites.
S. Boyd, N. Parikh, E. Chu, B. Peleato, J. Eckstein, et al
2011
Earlier work this paper cites.
M. Rubenstein, C. Ahler, and R. Nagpal, “Kilobot: A low cost scalable robot system for collective behaviors,” in Robotics and Automation (ICRA), 2012 IEEE International Conference on
2012
Earlier work this paper cites.
2013
Cited alongside, same era.
R. A. Knepper, T. Layton, J. Romanishin, and D. Rus, “Ikeabot: An autonomous multi-robot coordinated furniture assembly system,” in Robotics and Automation (ICRA), 2013 IEEE International Conference on
2013
Cited alongside, same era.
K. A. Potter, H. Arthur Woods, and S. Pincebourde, “Microclimatic challenges in global change biology,” Global change biology
2013
Cited alongside, same era.
2015
Cited alongside, same era.
2017
Later among the works it cites.
J. Stephan, J. Fink, V. Kumar, and A. Ribeiro, “Concurrent control of mobility and communication in multirobot systems,” IEEE Transactions on Robotics
2017
Later among the works it cites.
2017
Later among the works it cites.
R. Lowe, Y. Wu, A. Tamar, J. Harb, O. P. Abbeel, and I. Mordatch, “Multi-agent actor-critic for mixed cooperative-competitive environments,” in Advances in Neural Information Processing Systems
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in International Conference on Machine Learning
2015
Cited alongside, same era.
Software available from tensorflow.org
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015 · 2015
Cited alongside, same era.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in International Conference on Machine Learning
2016
Cited alongside, same era.
K. Solovey and D. Halperin, “On the hardness of unlabeled multi-robot motion planning,” The International Journal of Robotics Research
2016
Cited alongside, same era.
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel, “Benchmarking deep reinforcement learning for continuous control,” in International Conference on Machine Learning
2016
Cited alongside, same era.
2017
Cited alongside, same era.
A. Tampuu, T. Matiisen, D. Kodelja, I. Kuzovkin, K. Korjus, J. Aru, J. Aru, and R. Vicente, “Multiagent cooperation and competition with deep reinforcement learning,” PloS one
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
A. Khan, C. Zhang, N. Atanasov, K. Karydis, V. Kumar, and D. D. Lee, “Memory augmented control networks,” in International Conference on Learning Representations
2018
Closest in time.
J. Martens, J. Ba, and M. Johnson, “Kronecker-factored curvature approximations for recurrent neural networks,” in International Conference on Learning Representations
2018
Closest in time.