Fetching the paper…
Reading the bibliography…
We define a second-order neural network stochastic gradient training algorithm whose block-diagonal structure effectively amounts to normalizing the unit activations.
Improving the convergence of back-propagation learning with second-order methods
S. Becker and Y. LeCun · 1989
Earlier work this paper cites.
Efficient backprop
Y. Le Cun, L. Bottou, G. B. Orr, and K.-R. Müller · 1998
Earlier work this paper cites.
Efficient backprop
Yann Le Cun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 1998
Earlier work this paper cites.
Methods of Information Geometry
Sun-Ichi Amari and Hiroshi Nagaoka · 2000
Earlier work this paper cites.
Adaptive natural gradient learning algorithms for various stochastic models
Hyeyoung Park, Sun-ichi Amari, and Kenji Fukumizu · 2000
Earlier work this paper cites.
Numerical Optimization
Jorge Nocedal and Stephen J. Wright · 2006
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Centering neural network gradient factors
Nicol N. Schraudolph · 2012
Cited alongside, same era.
Lecture 6.5. RMSPROP: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Cited alongside, same era.
Improving deep neural networks for LVCSR using rectified linear units and dropout
George E. Dahl, Tara N. Sainath, and Geoffrey E. Hinton · 2013
Cited alongside, same era.
Riemannian metrics for neural networks
Yann Ollivier · 2013
Cited alongside, same era.
New insights and perspectives on the natural gradient method
James Martens · 2014
Cited alongside, same era.
Natural neural networks
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Later among the works it cites.
Optimizing neural networks with Kronecker-factored approximate curvature
James Martens and Roger B. Grosse · 2015
Later among the works it cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael I. Jordan, and Philipp Moritz · 2015
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2016
Later among the works it cites.
Practical Riemannian neural networks
Gaétan Marceau-Caron and Yann Ollivier · 2016
Later among the works it cites.
Understanding deep learning requires rethinking generalization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guillaume Desjardins, Karen Simonyan, Razvan Pascanu, and Koray Kavukcuoglu · 2015
Cited alongside, same era.
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Closest in time.