Fetching the paper…
Reading the bibliography…
Recurrent neural networks (RNNs) are notoriously difficult to train.
Linear Algebra
Hoffman, Kenneth and Kunze, Ray · 1971
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen Netzen. Diploma thesis, T.U. Münich, 1991
Hochreiter, S · 1991
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Yoshua, Simard, Patrice, and Frasconi, Paolo · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Yann, Bottou, Léon, Bengio, Yoshua, and Haffner, Patrick · 1998
Earlier work this paper cites.
Complex-valued neural networks: theories and applications , volume 5
Hirose, Akira · 2003
Earlier work this paper cites.
Theano: a CPU and GPU math expression compiler
Bergstra, James, Breuleux, Olivier, Bastien, Frédéric, Lamblin, Pascal, Pascanu, Razvan, Desjardins, Guillaume, Turian, Joseph, Warde-Farley, David, and Bengio, Yoshua · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, Xavier and Bengio, Yoshua · 2010
Earlier work this paper cites.
Fastfood - approximating kernel expansions in loglinear time
Le, Quoc, Sarlós, Tamás, and Smola, Alex · 2010
Cited alongside, same era.
Rectified linear units improve restricted boltzmann machines
Nair, Vinod and Hinton, Geoffrey E · 2010
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Pascanu, Razvan, Mikolov, Tomas, and Bengio, Yoshua · 2010
Cited alongside, same era.
Natural language processing (almost) from scratch
Collobert, Ronan, Weston, Jason, Bottou, Léon, Karlen, Michael, Kavukcuoglu, Koray, and Kuksa, Pavel · 2011
Cited alongside, same era.
Deep sparse rectifier neural networks
Glorot, Xavier, Bordes, Antoine, and Bengio, Yoshua · 2011
Cited alongside, same era.
Learning recurrent neural networks with hessian-free optimization
Martens, James and Sutskever, Ilya · 2011
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Later among the works it cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, Tijmen and Hinton, Geoffrey · 2012
Later among the works it cites.
On the properties of neural machine translation: Encoder–Decoder approaches
Cho, Kyunghyun, van Merriënboer, Bart, Bahdanau, Dzmitry, and Bengio, Yoshua · 2014
Later among the works it cites.
Graves, Alex, Wayne, Greg, and Danihelka, Ivo · 2014
Later among the works it cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, Andrew M., McLelland, James L., and Ganguli, Surya · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Comparison of the complex valued and real valued neural networks trained with gradient descent and random search algorithms
Zimmermann, Hans-Georg, Minin, Alexey, and Kusherbaeva, Victoria · 2011
Cited alongside, same era.
Deep neural networks for acoustic modeling in speech recognition
Hinton, Geoffrey, Deng, Li, Yu, Dong, Dahl, George, Mohamed, Abdel-rahman, Jaitly, Navdeep, Senior, Andrew, Vanhoucke, Vincent, Nguyen, Patrick, Sainath, Tara, and Kingsbury, Brian · 2012
Cited alongside, same era.
Le, Quoc V., Navdeep, Jaitly, and Hinton, Geoffrey E · 2015
Closest in time.
Deep fried convnets
Yang, Zichao, Moczulski, Marcin, Denil, Misha, de Freitas, Nando, Smola, Alex, Song, Le, and Wang, Ziyu · 2015
Closest in time.