Fetching the paper…
Reading the bibliography…
In recent years there have been many successes of using deep representations in reinforcement learning.
Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position
Fukushima, K · 1980
Earlier work this paper cites.
Advantage updating
Baird, L.C · 1993
Earlier work this paper cites.
Reinforcement learning for robots using neural networks
Lin, L.J · 1993
Earlier work this paper cites.
Advantage updating applied to a differential game
Harmon, M.E., Baird, L.C., and Klopf, A.H · 1995
Earlier work this paper cites.
Multi-player residual advantage learning with general function approximation
Harmon, M.E. and Baird, L.C · 1996
Earlier work this paper cites.
Introduction to reinforcement learning
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., Mcallester, D., Singh, S., and Mansour, Y · 2000
Earlier work this paper cites.
A theoretical and empirical analysis of Expected Sarsa
van Seijen, H., van Hasselt, H., Whiteson, S., and Wiering, M · 2009
Earlier work this paper cites.
Double Q-learning
van Hasselt, H · 2010
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Cited alongside, same era.
Advances in optimizing recurrent networks
Bengio, Y., Boulanger-Lewandowski, N., and Pascanu, R · 2013
Cited alongside, same era.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2013
Cited alongside, same era.
Deep learning for real-time Atari game play using offline Monte-Carlo tree search planning
Guo, X., Singh, S., Lee, H., Lewis, R. L., and Wang, X · 2014
Cited alongside, same era.
Multiple object recognition with visual attention
Ba, J., Mnih, V., and Kavukcuoglu, K · 2015
Cited alongside, same era.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Massively parallel methods for deep reinforcement learning
Nair, A., Srinivasan, P., Blackwell, S., Alcicek, C., Fearon, R., Maria, A. De, Panneershelvam, V., Suleyman, M., Beattie, C., Petersen, S., Legg, S., Mnih, V., Kavukcuoglu, K., and Silver, D · 2015
Closest in time.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M. I., and Abbeel, P · 2015
Closest in time.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S., and Abbeel, P · 2015
Closest in time.
Deep reinforcement learning with double Q-learning
van Hasselt, H., Guez, A., and Silver, D · 2015
Closest in time.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J. T., Boedecker, J., and Riedmiller, M. A · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2015
Cited alongside, same era.
Move Evaluation in Go Using Deep Convolutional Neural Networks
Maddison, C. J., Huang, A., Sutskever, I., and Silver, D · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Cited alongside, same era.
Closest in time.
Increasing the action gap: New operators for reinforcement learning
Bellemare, M. G., Ostrovski, G., Guez, A., Thomas, P. S., and Munos, R · 2016
Closest in time.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Closest in time.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C.J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Closest in time.