End-to-end training of deep visuomotor policies
Original
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2015
Later among the works it cites.
Recurrent reinforcement learning: A hybrid approach
Li, X., Li, L., Gao, J., He, X., Chen, J., Deng, L., and He, J · 2015
Later among the works it cites.
Move Evaluation in Go Using Deep Convolutional Neural Networks
Maddison, C. J., Huang, A., Sutskever, I., and Silver, D · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Later among the works it cites.
Massively parallel methods for deep reinforcement learning
Nair, A., Srinivasan, P., Blackwell, S., Alcicek, C., Fearon, R., Maria, A. De, Panneershelvam, V., Suleyman, M., Beattie, C., Petersen, S., Legg, S., Mnih, V., Kavukcuoglu, K., and Silver, D · 2015
Later among the works it cites.
Language understanding for text-based games using deep reinforcement learning
Narasimhan, K., Kulkarni, T., and Barzilay, R · 2015
Later among the works it cites.
Action-conditional video prediction using deep networks in Atari games
Oh, J., Guo, X., Lee, H., Lewis, R. L., and Singh, S · 2015
Later among the works it cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Original
Stadie, B. C., Levine, S., and Abbeel, P · 2015
Later among the works it cites.
Multiagent cooperation and competition with deep reinforcement learning
Original
Tampuu, A., Matiisen, T., Kodelja, D., Kuzovkin, I., Korjus, K., Aru, J., Aru, J., and Vicente, R · 2015
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., de Freitas, N., and Lanctot, M · 2015
Later among the works it cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J. T., Boedecker, J., and Riedmiller, M. A · 2015
Later among the works it cites.
Increasing the action gap: New operators for reinforcement learning
Bellemare, M. G., Ostrovski, G., Guez, A., Thomas, P. S., and Munos, R · 2016
Closest in time.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Closest in time.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C.J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Closest in time.
Deep reinforcement learning with double Q-learning
van Hasselt, H., Guez, A., and Silver, D · 2016
Closest in time.