Fetching the paper…
Reading the bibliography…
Natural gradient descent is an optimization method traditionally motivated from the perspective of information geometry, and works well for many applications as an alternative to stochastic gradient descent.
On-line learning for very large data sets
L. Bottou and Y. LeCun · 1904
Earlier work this paper cites.
On the stability of inverse problems
A. N. Tikhonov · 1943
Earlier work this paper cites.
A method for the solution of certain non-linear problems in least squares
K. Levenberg · 1944
Earlier work this paper cites.
An algorithm for least-squares estimation of nonlinear parameters
D. W. Marquardt · 1963
Earlier work this paper cites.
Theory of adaptive pattern classifiers
S. Amari · 1967
Earlier work this paper cites.
The Levenberg-Marquardt algorithm: implementation and theory
J. Moré · 1978
Earlier work this paper cites.
Inexact newton methods
R. S. Dembo, S. C. Eisenstat, and T. Steihaug · 1982
Earlier work this paper cites.
Computing a trust region step
J. J. Moré and D. C. Sorensen · 1983
Earlier work this paper cites.
The conjugate gradient method and trust regions in large scale optimization
T. Steihaug · 1983
Earlier work this paper cites.
Preconditioning of truncated-newton methods
S. G. Nash · 1985
Earlier work this paper cites.
Trace bounds on the solution of the algebraic matrix Riccati and Lyapunov equation
S.-D. Wang, T.-S. Kuo, and C.-F. Hsu · 1986
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
S. Becker and Y. LeCun · 1989
Earlier work this paper cites.
Training multilayer perceptrons with the extended kalman algorithm
S. Singhal and L. Wu · 1989
Earlier work this paper cites.
Euler’s constant
R. M. Young · 1991
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Comparative analysis of backpropagation and the extended kalman filter for training multilayer perceptrons
D. W. Ruck, S. K. Rogers, M. Kabrisky, P. S. Maybeck, and M. E. Oxley · 1992
Earlier work this paper cites.
Tangent prop-a formalism for specifying selected invariances in an adaptive network
P. Simard, B. Victorri, Y. LeCun, and J. Denker · 1992
Earlier work this paper cites.
Numerical methods for unconstrained optimization and nonlinear equations , volume 16
J. E. Dennis Jr and R. B. Schnabel · 1996
Earlier work this paper cites.
Natural gradient works efficiently in learning
S.-I. Amari · 1998
Earlier work this paper cites.
Adaptive blind signal processing-neural network approaches
S.-i. Amari and A. Cichocki · 1998
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, G. Orr, and K. Müller · 1998
Earlier work this paper cites.
A statistical study of on-line learning
N. Murata · 1998
Earlier work this paper cites.
Methods of Information Geometry , volume 191 of Translations of Mathematical monographs
S. Amari and H. Nagaoka · 2000
Earlier work this paper cites.
Trust region methods
A. R. Conn, N. I. Gould, and P. L. Toint · 2000
Earlier work this paper cites.
On “natural” learning and pruning in multilayered perceptrons
T. Heskes · 2000
Earlier work this paper cites.
Gradient Flow in Recurrent Nets: the Difficulty of Learning Long-Term Dependencies
S. Hochreiter, F. F. Informatik, Y. Bengio, P. Frasconi, and J. Schmidhuber · 2000
Cited alongside, same era.
Iterative solution of nonlinear equations in several variables
J. M. Ortega and W. C. Rheinboldt · 2000
Cited alongside, same era.
Adaptive natural gradient learning algorithms for various stochastic models
H. Park, S.-I. Amari, and K. Fukumizu · 2000
Cited alongside, same era.
The ubiquitous kronecker product
C. F. Van Loan · 2000
Cited alongside, same era.
Fast curvature matrix-vector products for second-order gradient descent
N. N. Schraudolph · 2002
Cited alongside, same era.
Stochastic approximation and recursive algorithms and applications , volume 35
H. Kushner and G. G. Yin · 2003
Cited alongside, same era.
Practical methods of optimization
R. Fletcher · 2013
Later among the works it cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Later among the works it cites.
Introductory lectures on convex optimization: A basic course , volume 87
Y. Nesterov · 2013
Later among the works it cites.
No more pesky learning rates
T. Schaul, S. Zhang, and Y. LeCun · 2013
Later among the works it cites.
ADADELTA: An adaptive learning rate method
M. D. Zeiler · 2013
Later among the works it cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio · 2014
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Numerical optimization
J. Nocedal and S. J. Wright · 2006
Cited alongside, same era.
Methods of information geometry , volume 191
S.-i. Amari and H. Nagaoka · 2007
Cited alongside, same era.
Logarithmic regret algorithms for online convex optimization
E. Hazan, A. Agarwal, and S. Kale · 2007
Cited alongside, same era.
Topmoumoute online natural gradient algorithm
N. Le Roux, P.-a. Manzagol, and Y. Bengio · 2008
Cited alongside, same era.
Natural actor-critic
J. Peters and S. Schaal · 2008
Cited alongside, same era.
SGD-QN: Careful Quasi-Newton Stochastic Gradient Descent
A. Bordes, L. Bottou, and P. Gallinari · 2009
Cited alongside, same era.
Constant step size least-mean-square: Bias-variance trade-offs and optimal sampling distributions
A. Défossez and F. Bach · 2014
Closest in time.
Competing with the empirical risk minimizer in a single pass
R. Frostig, R. Ge, S. M. Kakade, and A. Sidford · 2014
Closest in time.
Revisiting natural gradient for deep networks
R. Pascanu and Y. Bengio · 2014
Closest in time.
Adam: A method for stochastic optimization
J. Ba and D. Kingma · 2015
Closest in time.
Natural neural networks
G. Desjardins, K. Simonyan, R. Pascanu, and K. Kavukcuoglu · 2015
Closest in time.
From averaging to acceleration, there is only a step-size
N. Flammarion and F. Bach · 2015
Closest in time.
Scaling up natural gradient by sparsely factorizing the inverse fisher matrix
R. Grosse and R. Salakhudinov · 2015
Closest in time.
Optimizing neural networks with Kronecker-factored approximate curvature
J. Martens and R. Grosse · 2015
Closest in time.
Riemannian metrics for neural networks i: feedforward networks
Y. Ollivier · 2015
Closest in time.
Second-order optimization for neural networks
J. Martens · 2016
Closest in time.
Practical gauss-newton optimisation for deep learning
A. Botev, H. Ritter, and D. Barber · 2017
Closest in time.
N. Loizou and P. Richtárik · 2017
Closest in time.
Gradient descent finds global minima of deep neural networks
S. S. Du, J. D. Lee, H. Li, L. Wang, and X. Zhai · 2018
Closest in time.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Closest in time.
Online natural gradient as a kalman filter
Y. Ollivier et al · 2018
Closest in time.
A gram-gauss-newton method learning overparameterized deep neural networks for regression problems
T. Cai, R. Gao, J. Hou, S. Chen, D. Wang, D. He, Z. Zhang, and L. Wang · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
J. Lee, L. Xiao, S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington · 2019
Closest in time.
Fast convergence of natural gradient descent for over-parameterized neural networks
G. Zhang, J. Martens, and R. B. Grosse · 2019
Closest in time.