Fetching the paper…
Reading the bibliography…
Second-order optimization methods have the ability to accelerate convergence by modifying the gradient through the curvature matrix.
Symmetry, 0-1 matrices and jacobians: A review
Jan R Magnus and Heinz Neudecker · 1986
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second-order methods
S. Becker and Y. Lecun · 1988
Earlier work this paper cites.
Stochastic gradient learning in neural networks
Léon Bottou · 1991
Earlier work this paper cites.
Kronecker products and the vec and vech operators
David A Harville · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y Lecun and L Bottou · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
Training the random neural network using quasi-Newton methods
Aristidis Likas and Andreas Stafylopatis · 2000
Earlier work this paper cites.
A well-conditioned estimator for large-dimensional covariance matrices
Olivier Ledoit and Michael Wolf · 2004
Earlier work this paper cites.
Maximum-margin matrix factorization
Nathan Srebro, Jason Rennie, and Tommi S Jaakkola · 2005
Earlier work this paper cites.
Methods of Information Geometry , volume 191
Shun-ichi Amari and Hiroshi Nagaoka · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Deep learning via Hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
Benjamin Recht, Maryam Fazel, and Pablo A. Parrilo · 2010
Earlier work this paper cites.
High-dimensional covariance matrix estimation in approximate factor models
Jianqing Fan and Liao Martina Mincheva · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi John, Hazan Elad, and Singer Yoram · 2011
Cited alongside, same era.
Deep learning via Hessian-free optimization
James Martens and Sutskever Ilya · 2011
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Cited alongside, same era.
Training neural networks with stochastic Hessian-free optimization
Ryan Kiros · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Accelerating Hessian-free Gauss-Newton full-waveform inversion via l-BFGS preconditioned conjugate-gradient algorithm
Wenyong Pan, Kristopher A Innanen, and Wenyuan Liao · 2017
Later among the works it cites.
Eigenvalue corrected noisy natural gradient
Juhan Bae, Guodong Zhang, and Roger Grosse · 2018
Later among the works it cites.
The properties of partial trace and block trace operators of partitioned matrices
Katarzyna Filipiak, Daniel Klein, and Erika Vojtková · 2018
Later among the works it cites.
Fast approximate natural gradient descent in a kronecker factored eigenbasis
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas, and Pascal Vincent · 2018
Later among the works it cites.
Kronecker-factored curvature approximations for recurrent neural networks
James Martens, Jimmy Ba, and Matt Johnson · 2018
Later among the works it cites.
Noisy natural gradient as variational inference
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
New insights and perspectives on the natural gradient method
James Martens · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
A kronecker-factored approximate fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Equivariant and scale-free tucker decomposition models
Peter D. Hoff · 2016
Cited alongside, same era.
Guodong Zhang, Shengyang Sun, David Duvenaud, and Roger Grosse · 2018
Later among the works it cites.
Quasi-Newton methods for deep learning: Forget the past, just sample
Albert S Berahas, Majid Jahani, and Martin Takáč · 2019
Later among the works it cites.
Pathological spectra of the fisher information metric and its variants in deep neural networks
Ryo Karakida, Shotaro Akaho, and Shun-ichi Amari · 2019
Later among the works it cites.
Estimation of the kronecker covariance model by partial means and quadratic form
Oliver B Linton and Haihan Tang · 2019
Later among the works it cites.
Large-scale distributed second-order optimization using kronecker-factored approximate curvature for deep convolutional neural networks
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka · 2019
Later among the works it cites.
Three mechanisms of weight decay regularization
Guodong Zhang, Chaoqi Wang, Bowen Xu, and Roger Grosse · 2019
Later among the works it cites.
Practical quasi-Newton methods for training deep neural networks
Donald Goldfarb, Yi Ren, and Achraf Bahamou · 2020
Closest in time.
Convolutional neural network training with distributed K-FAC
J. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu, and Ian T. Foster · 2020
Closest in time.
ADAHESSIAN: An adaptive second order optimizer for machine learning
Zhewei Yao, Amir Gholami, Sheng Shen, Kurt Keutzer, and Michael W Mahoney · 2020
Closest in time.