Fetching the paper…
Reading the bibliography…
Training very deep networks is an important open problem in machine learning.
Untersuchungen zu dynamischen neuronalen Netzen. Diploma thesis, T.U. Münich, 1991
Hochreiter, S · 1991
Earlier work this paper cites.
The problem of learning long-term dependencies in recurrent networks
Bengio, Y., Frasconi, P., and Simard, P · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., and Frasconi, P · 1994
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies, 2001
Hochreiter, Sepp, Bengio, Yoshua, Frasconi, Paolo, and Schmidhuber, Jürgen · 2001
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Hinton, Geoffrey E and Salakhutdinov, Ruslan R · 2006
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
Glorot, Xavier and Bengio, Yoshua · 2010
Cited alongside, same era.
Deep learning via hessian-free optimization
Martens, James · 2010
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Pascanu, Razvan, Mikolov, Tomas, and Bengio, Yoshua · 2012
Later among the works it cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, Andrew M, McClelland, James L, and Ganguli, Surya · 2013
Later among the works it cites.
On the importance of initialization and momentum in deep learning
Sutskever, Ilya, Martens, James, Dahl, George, and Hinton, Geoffrey · 2013
Later among the works it cites.
On the saddle point problem for non-convex optimization
Pascanu, Razvan, Dauphin, Yann N, Ganguli, Surya, and Bengio, Yoshua · 2014
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…