Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., … and Kavukcuoglu, K. (2016, June). Asynchronous methods for deep reinforcement learning. In International Conference on Machine Learning (pp. 1928-1937)
1937
Earlier work this paper cites.
Bellman, R. (1952). On the theory of dynamic programming. Proceedings of the National Academy of Sciences , 38(8), 716-719
1952
Earlier work this paper cites.
Minsky, M. L. (1954). Theory of neural-analog reinforcement systems and its application to the brain model problem. Princeton University
1954
Earlier work this paper cites.
Klopf, A. (1972). Brain function and adaptive systems: a heterostatic theory. Air Force Cambridge Research Laboratories Special Report
1972
Earlier work this paper cites.
Mahadevan, S., and Connell, J. (1992). Automatic programming of behavior-based robots using reinforcement learning. Artificial Intelligence , 55(2-3), 311-365
1992
Earlier work this paper cites.
Watkins, C. J., and Dayan, P. (1992). Q-learning. Machine Learning , 8(3-4), 279-292
1992
Earlier work this paper cites.
Tan, M. (1993). Multi-agent reinforcement learning: independent vs. cooperative agents. In Proceedings of the Tenth International Conference on Machine Learning (pp. 330-337)
1993
Earlier work this paper cites.
Benbrahim, H., and Franklin, J. A. (1997). Biped dynamic walking using reinforcement learning. Robotics and Autonomous Systems , 22(3-4), 283-302
1997
Earlier work this paper cites.
Schaal, S. (1997). Learning from demonstration. In Advances in Neural Information Processing Systems (pp. 1040-1046)
1997
Earlier work this paper cites.
Tsitsiklis, J. N., and Van Roy, B. (1997). Analysis of temporal-diffference learning with function approximation. In Advances in Neural Information Processing Systems (pp. 1075-1081)
1997
Earlier work this paper cites.
Sutton, R. S., and Barto, A. G. (1998). Reinforcement learning: An introduction. MIT press
1998
Earlier work this paper cites.
Konda, V. R., and Tsitsiklis, J. N. (2000). Actor-critic algorithms. In Advances in Neural Information Processing Systems (pp. 1008-1014)
2000
Earlier work this paper cites.
Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N. (2016, June). Dueling network architectures for deep reinforcement learning. In International Conference on Machine Learning (pp. 1995-2003)
2003
Earlier work this paper cites.
de Cote, E. M., Lazaric, A., and Restelli, M. (2006, May). Learning to cooperate in multi-agent social dilemmas. In Proceedings of the 5th International Joint Conference on Autonomous Agents and Multiagent Systems (pp. 783-785). ACM
2006
Earlier work this paper cites.
Zinkevich, M., Greenwald, A., and Littman, M. L. (2006). Cyclic equilibria in Markov games. In Advances in Neural Information Processing Systems (pp. 1641-1648)
2006
Earlier work this paper cites.
Harati, A., Ahmadabadi, M. N., and Araabi, B. N. (2007). Knowledge-based multiagent credit assignment: a study on task type and critic information. IEEE Systems Journal , 1(1), 55-67
2007
Earlier work this paper cites.
Matignon, L., Laurent, G., and Le Fort-Piat, N. (2007, October). Hysteretic Q-Learning: an algorithm for decentralized reinforcement learning in cooperative multi-agent teams. In IEEE/RSJ International Conference on Intelligent Robots and Systems (pp. 64-69)
2007
Earlier work this paper cites.
Shoham, Y., Powers, R., and Grenager, T. (2007). If multi-agent learning is the answer, what is the question?. Artificial Intelligence , 171(7), 365-377
2007
Earlier work this paper cites.
Busoniu, L., Babuska, R., and De Schutter, B. (2008). A comprehensive survey of multiagent reinforcement learning. IEEE Transactions on Systems, Man, and Cybernetics-Part C: Applications and Reviews , 38 (2), 156-172
2008
Earlier work this paper cites.
Riedmiller, M., Gabel, T., Hafner, R., and Lange, S. (2009). Reinforcement learning for robot soccer. Autonomous Robots , 27(1), 55-73
2009
Earlier work this paper cites.
Hasselt, H. V. (2010). Double Q-learning. In Advances in Neural Information Processing Systems (pp. 2613-2621)
2010
Earlier work this paper cites.
Janssen, M. A., Holahan, R., Lee, A., and Ostrom, E. (2010). Lab experiments for the study of social-ecological systems. Science , 328(5978), 613-617
2010
Earlier work this paper cites.
Brandouy, O., Mathieu, P., and Veryzhenko, I. (2011, January). On the design of agent-based artificial stock markets. In International Conference on Agents and Artificial Intelligence (pp. 350-364)
2011
Earlier work this paper cites.
Chung, T. H., Hollinger, G. A., and Isler, V. (2011). Search and pursuit-evasion in mobile robotics. Autonomous Robots , 31(4), 299
2011
Earlier work this paper cites.
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012). Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems (pp. 1097-1105)
2012
Earlier work this paper cites.
Oliehoek, F. A. (2012). Decentralized POMDPs. In Reinforcement Learning (pp. 471-503). Springer, Berlin, Heidelberg
2012
Earlier work this paper cites.
Abdallah, S., and Kaisers, M. (2013, May). Addressing the policy-bias of Q-learning by repeating updates. In Proceedings of the 12th International Conference on Autonomous Agents and Multiagent Systems (pp. 1045-1052)
2013
Earlier work this paper cites.
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013). The arcade learning environment: an evaluation platform for general agents. Journal of Artificial Intelligence Research , 47, 253-279
2013
Earlier work this paper cites.
Mulling, K., Kober, J., Kroemer, O., and Peters, J. (2013). Learning to select and generalize striking movements in robot table tennis. The International Journal of Robotics Research , 32(3), 263-279
2013
Earlier work this paper cites.
Deng, L., and Yu, D. (2014). Deep learning: methods and applications. Foundations and Trends in Signal Processing , 7(3–4), 197-387
2014
Earlier work this paper cites.
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014, January). Deterministic policy gradient algorithms. In International Conference on Machine Learning (pp. 387-395)
2014
Earlier work this paper cites.
Bloembergen, D., Tuyls, K., Hennes, D., and Kaisers, M. (2015). Evolutionary dynamics of multi-agent learning: a survey. Journal of Artificial Intelligence Research , 53, 659-697
2015
Earlier work this paper cites.
Fernandez-Gauna, B., Etxeberria-Agiriano, I., and Grana, M. (2015). Learning multirobot hose transportation and deployment by distributed round-robin Q-learning. PloS One , 10(7), e0127129
2015
Earlier work this paper cites.
Hausknecht, M., and Stone, P. (2015). Deep recurrent Q-learning for partially observable MDPs. CoRR, abs/1507.06527 , 7(1)
Original
2015
Earlier work this paper cites.
He, J., Peng, J., Jiang, F., Qin, G., and Liu, W. (2015). A distributed Q learning spectrum decision scheme for cognitive radio sensor network. International Journal of Distributed Sensor Networks , 11(5), 301317
2015
Earlier work this paper cites.
Heess, N., Hunt, J. J., Lillicrap, T. P., and Silver, D. (2015). Memory-based control with recurrent neural networks. arXiv preprint arXiv:1512.04455
Original
2015
Earlier work this paper cites.
Hussin, M., Hamid, N. A. W. A., and Kasmiran, K. A. (2015). Improving reliability in resource management through adaptive reinforcement learning for distributed systems. Journal of Parallel and Distributed Computing , 75, 93-100
2015
Earlier work this paper cites.
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., … and Wierstra, D. (2015). Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971
Original
2015
Earlier work this paper cites.
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., … and Petersen, S. (2015). Human-level control through deep reinforcement learning. Nature , 518(7540), 529-533
2015
Earlier work this paper cites.
Parisotto, E., Ba, J. L., and Salakhutdinov, R. (2015). Actor-mimic: deep multitask and transfer reinforcement learning. arXiv preprint arXiv:1511.06342
Original
2015
Earlier work this paper cites.
Rahman, M. S., Mahmud, M. A., Pota, H. R., Hossain, M. J., and Orchi, T. F. (2015). Distributed multi-agent-based protection scheme for transient stability enhancement in power systems. International Journal of Emerging Electric Power Systems , 16(2), 117-129
2015
Earlier work this paper cites.