Fetching the paper…
Reading the bibliography…
We describe four algorithms for neural network training, each adapted to different scalability constraints.
Riemannian geometry
S. Gallot, D. Hulin, and J. Lafontaine · 1987
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
Sue Becker and Yann LeCun · 1988
Earlier work this paper cites.
Methods of information geometry
Shun-ichi Amari and Hiroshi Nagaoka · 1993
Earlier work this paper cites.
Iterative weighted least squares algorithms for neural networks classifiers
Takio Kurita · 1994
Earlier work this paper cites.
Efficient backprop
Yann LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 1996
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-ichi Amari · 1998
Earlier work this paper cites.
Adaptive method of realizing natural gradient learning for multilayer perceptrons
Shun-ichi Amari, Hyeyoung Park, and Kenji Fukumizu · 2000
Cited alongside, same era.
Artificial Intelligence: A Modern Approach
Stuart Russell and Peter Norvig · 2003
Cited alongside, same era.
Gradient flows in metric spaces and in the space of probability measures
Luigi Ambrosio, Nicola Gigli, and Giuseppe Savaré · 2005
Cited alongside, same era.
Pattern recognition and machine learning
Christopher M. Bishop · 2006
Cited alongside, same era.
Topmoumoute online natural gradient algorithm
Nicolas Le Roux, Pierre-Antoine Manzagol, and Yoshua Bengio · 2007
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John C. Duchi, Elad Hazan, and Yoram Singer · 2011
Later among the works it cites.
On optimization methods for deep learning
Quoc V. Le, Jiquan Ngiam, Adam Coates, Ahbik Lahiri, Bobby Prochnow, and Andrew Y. Ng · 2011
Later among the works it cites.
Information-Geometric Optimization algorithms: A unifying picture via invariance principles
Yann Ollivier, Ludovic Arnold, Anne Auger, and Nikolaus Hansen · 2011
Later among the works it cites.
Riemannian metrics for neural networks II: recurrent networks and learning symbolic data sequences
Yann Ollivier · 2013
Closest in time.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio · 2013
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Closest in time.