Connectionist learning procedures
Hinton, G. E · 1989
Earlier work this paper cites.
Learning invariance from transformation sequences
Földiák, P · 1991
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. and Juditsky, A · 1992
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Koehn, P., Hoang, H., Birch, A., Callison-Burch, C., Federico, M., Bertoldi, N., Cowan, B., Shen, W., Moran, C., Zens, R., Dyer, C., Bojar, O., Constantin, A., and Herbst, E · 2007
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Cernocký, J., and Khudanpur, S · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Low complexity proto-value function learning from sensory observations with incremental slow feature analysis
Luciw, M. and Schmidhuber, J · 2012
Earlier work this paper cites.
Context dependent recurrent neural network language model
Mikolov, T. and Zweig, G · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Wan, L., Zeiler, M., Zhang, S., LeCun, Y, and Fergus, R · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Original
Kingma, D. and Ba, J · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Original
Zaremba, W., Sutskever, I., and Vinyals, O · 2014
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Original
Hardt, M., Recht, B., and Singer, Y · 2015
Earlier work this paper cites.