Fetching the paper…
Reading the bibliography…
Most neural networks are trained using first-order optimization methods, which are sensitive to the parameterization of the model.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Methods of Information Geometry
S. Amari and H. Nagaoka · 2000
Earlier work this paper cites.
Adaptive method of realizing natural gradient learning for multilayer perceptrons
S. Amari, H. Park, and K. Fukumizu · 2000
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Nicol N Schraudolph · 2002
Earlier work this paper cites.
Smooth manifolds
John M Lee · 2003
Earlier work this paper cites.
The tradeoffs of large scale learning
L. Bottou and O. Bousquet · 2007
Earlier work this paper cites.
Topmoumoute online natural gradient algorithm
N. Le Roux, P.-A. Manzagol, and Y. Bengio · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Deep learning via hessian-free optimization
J. Martens · 2010
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Enhanced gradient and adaptive learning rate for training restricted Boltzmann machines
K. Cho, T. Raiko, and A. Ilin · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Deep Boltzmann machines and the centering trick
G. Montavon and K.-R. Müller · 2012
Cited alongside, same era.
Deep learning made easier by linear transformations in perceptrons
T. Raiko, H. Valpola, and Y. LeCun · 2012
Cited alongside, same era.
Krylov subspace descent for deep learning
Oriol Vinyals and Daniel Povey · 2012
Cited alongside, same era.
Stochastic gradient descent on riemannian manifolds
Silvere Bonnabel et al · 2013
Cited alongside, same era.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio · 2013
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger B Grosse · 2015
Later among the works it cites.
Riemannian metrics for neural networks I: feedforward networks
Y. Ollivier · 2015
Later among the works it cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Later among the works it cites.
A Kronecker-factored approximate Fisher matrix for convolution layers
R. Grosse and J. Martens · 2016
Later among the works it cites.
Distributed second-order optimization using Kronecker-factored approximations
J. Ba, R. Grosse, and J. Martens · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
New insights and perspectives on the natural gradient method
James Martens · 2014
Cited alongside, same era.
Natural neural networks
G. Desjardins, K. Simonyan, and R. Pascanu · 2015
Cited alongside, same era.
Natural neural networks
Guillaume Desjardins, Karen Simonyan, Razvan Pascanu, et al · 2015
Cited alongside, same era.
Scaling up natural gradient by sparsely factorizing the inverse Fisher matrix
R. B. Grosse and R. Salakhutdinov · 2015
Cited alongside, same era.
Batch normalization: accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Adam: a method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Yuhuai Wu, Elman Mansimov, Shun Liao, Roger Grosse, and Jimmy Ba · 2017
Later among the works it cites.
Noisy natural gradient as variational inference
Guodong Zhang, Shengyang Sun, David Duvenaud, and Roger Grosse · 2017
Later among the works it cites.
Kronecker-factored curvature approximations for recurrent neural networks
James Martens, Jimmy Ba, and Matt Johnson · 2018
Closest in time.
Online structured laplace approximations for overcoming catastrophic forgetting
Hippolyt Ritter, Aleksandar Botev, and David Barber · 2018
Closest in time.
A scalable laplace approximation for neural networks
Hippolyt Ritter, Aleksandar Botev, and David Barber · 2018
Closest in time.
Accelerating natural gradient with higher-order invariance
Yang Song and Stefano Ermon · 2018
Closest in time.