Fetching the paper…
Reading the bibliography…
We present practical Levenberg-Marquardt variants of Gauss-Newton and natural gradient methods for solving non-convex optimization problems that arise in training deep neural networks involving enormous numbers of variables and huge data sets.
The levenberg-marquardt algorithm: implementation and theory, numerical analysis
J. J. Moré · 1977
Earlier work this paper cites.
Natural gradient works efficiently in learning
S.-I. Amari · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
N. Qian · 1999
Earlier work this paper cites.
Numerical optimization
S. Wright and J. Nocedal · 1999
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
N. N. Schraudolph · 2002
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Deep learning via hessian-free optimization
J. Martens · 2010
Cited alongside, same era.
LIBSVM: A library for support vector machines
C.-C. Chang and C.-J. Lin · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Krylov subspace descent for deep learning
O. Vinyals and D. Povey · 2012
Cited alongside, same era.
Training neural networks with stochastic hessian-free optimization
R. Kiros · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Practical Gauss-Newton optimisation for deep learning
A. Botev, H. Ritter, and D. Barber · 2017
Later among the works it cites.
Improving generalization performance by switching from adam to sgd
N. S. Keskar and R. Socher · 2017
Later among the works it cites.
Exact natural gradient in deep linear networks and its application to the nonlinear case
A. Bernacchia, M. Lengyel, and G. Hennequin · 2018
Later among the works it cites.
Adaptive sampling strategies for stochastic optimization
R. Bollapragada, R. Byrd, and J. Nocedal · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Optimizing neural networks with Kronecker-factored approximate curvature
J. Martens and R. Grosse · 2015
Cited alongside, same era.
T. Cai, R. Gao, J. Hou, S. Chen, D. Wang, D. He, Z. Zhang, and L. Wang · 2019
Closest in time.
Fast convergence of natural gradient descent for overparameterized neural networks
G. Zhang, J. Martens, and R. Grosse · 2019
Closest in time.