A markovian decision process
Bellman, R · 1957
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
LeCun, Y., Boser, B. E., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W. E., and Jackel, L. D · 1989
Earlier work this paper cites.
Self-organizing neural network that discovers surfaces in random-dot stereograms
Becker, S. and Hinton, G. E · 1992
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P · 2002
Earlier work this paper cites.
Best practices for convolutional neural networks applied to visual document analysis
Simard, P. Y., Steinkraus, D., Platt, J. C., et al · 2003
Earlier work this paper cites.
Invariant causal prediction for block mdps
Original
Zhang, A., Lyle, C., Sodhani, S., Filos, A., Kwiatkowska, M., Pineau, J., Gal, Y., and Precup, D · 2003
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 2004
Earlier work this paper cites.
Learning invariant representations for reinforcement learning without reconstruction
Original
Zhang, A., McAllister, R., Calandra, R., Gal, Y., and Levine, S · 2006
Earlier work this paper cites.
High-performance neural networks for visual object classification
Original
Ciresan, D. C., Meier, U., Masci, J., Gambardella, L. M., and Schmidhuber, J · 2011
Earlier work this paper cites.
High-performance neural networks for visual object classification
Original
Cireşan, D. C., Meier, U., Masci, J., Gambardella, L. M., and Schmidhuber, J · 2011
Earlier work this paper cites.
Multi-column deep neural networks for image classification
Ciregan, D., Meier, U., and Schmidhuber, J · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Original
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G. E., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Original
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Original
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., and Moritz, P · 2015
Earlier work this paper cites.
Discriminative unsupervised feature learning with exemplar convolutional neural networks
Dosovitskiy, A., Fischer, P., Springenberg, J. T., Riedmiller, M. A., and Brox, T · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Original
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2016
Earlier work this paper cites.
Improved regularization of convolutional neural networks with cutout
Original
DeVries, T. and Taylor, G. W · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.