Fetching the paper…
Reading the bibliography…
Parameter-specific adaptive learning rate methods are computationally efficient ways to reduce the ill-conditioning problems encountered when training large deep networks.
Condition numbers and equilibration of matrices
Sluis, AVD · 1969
Earlier work this paper cites.
A simple estimate of the condition number of a linear system
Guggenheimer, Heinrich W., Edelman, Alan S., and Johnson, Charles R · 1995
Earlier work this paper cites.
Efficient backprop
LeCun, Yann, Bottou, Léon, Orr, Genevieve B., and Müller, Klaus-Robert · 1998
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Schraudolph, Nicol N · 2002
Earlier work this paper cites.
An estimator for the diagonal of a matrix
Bekas, Costas, Kokiopoulou, Effrosyni, and Saad, Yousef · 2007
Earlier work this paper cites.
Numerical Linear Algebra and Applications, Second Edition
Datta, Biswa Nath · 2010
Earlier work this paper cites.
Deep learning via Hessian-free optimization
Martens, J · 2010
Earlier work this paper cites.
Matrix-free approximate equilibration
Bradley, Andrew M and Murray, Walter · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, and Singer, Yoram · 2011
Cited alongside, same era.
Krylov subspace descent for deep learning
Vinyals, Oriol and Povey, Daniel · 2011
Cited alongside, same era.
Theano: new features and speed improvements
Bastien, Frédéric, Lamblin, Pascal, Pascanu, Razvan, Bergstra, James, Goodfellow, Ian J., Bergeron, Arnaud, Bouchard, Nicolas, and Bengio, Yoshua · 2012
Cited alongside, same era.
Estimating the hessian by back-propagating curvature
Martens, James, Sutskever, Ilya, and Swersky, Kevin · 2012
Cited alongside, same era.
ADADELTA: an adaptive learning rate method
Zeiler, Matthew D · 2012
Later among the works it cites.
Unit tests for stochastic optimization
Schaul, Tom, Antonoglou, Ioannis, and Silver, David · 2013
Later among the works it cites.
On the importance of initialization and momentum in deep learning
Sutskever, Ilya, Martens, James, Dahl, George, and Hinton, Geoffrey · 2013
Later among the works it cites.
The loss surface of multilayer networks, 2014
Choromanska, Anna, Henaff, Mikael, Mathieu, Michael, Arous, Gérard Ben, and LeCun, Yann · 2014
Later among the works it cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Yann, Pascanu, Razvan, Gulcehre, Caglar, Cho, Kyunghyun, Ganguli, Surya, and Bengio, Yoshua · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, Tijmen and Hinton, Geoffrey · 2012
Cited alongside, same era.
Revisiting natural gradient for deep networks
Pascanu, Razvan and Bengio, Yoshua · 2014
Later among the works it cites.