Fetching the paper…
Reading the bibliography…
In the context of deep learning, many optimization methods use gradient covariance information in order to accelerate the convergence of Stochastic Gradient Descent.
Fast exact multiplication by the hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Nicol N Schraudolph · 2002
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
Comparison of two-level preconditioners derived from deflation, domain decomposition and multigrid methods
Jok Man Tang, Reinhard Nabben, Cornelis Vuik, and Yogi A Erlangga · 2009
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Adaptive cubic regularisation methods for unconstrained optimization. part ii: worst-case function-and derivative-evaluation complexity
Coralia Cartis, Nicholas IM Gould, and Philippe L Toint · 2011
Earlier work this paper cites.
Improved preconditioner for hessian free optimization
Olivier Chapelle and Dumitru Erhan · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Analysis of a two-level schwarz method with coarse spaces based on local dirichlet-to-neumann maps
Victorita Dolean, Frédéric Nataf, Robert Scheichl, and Nicole Spillane · 2012
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
An algebraic multigrid method with guaranteed convergence rate
Artem Napov and Yvan Notay · 2012
Earlier work this paper cites.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio · 2013
Cited alongside, same era.
Scalable domain decomposition preconditioners for heterogeneous elliptic problems
Pierre Jolivet, Frédéric Hecht, Frédéric Nataf, and Christophe Prud’Homme · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
New insights and perspectives on the natural gradient method
James Martens · 2014
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
A progressive batching l-bfgs method for machine learning
Raghu Bollapragada, Dheevatsa Mudigere, Jorge Nocedal, Hao-Jun Michael Shi, and Ping Tak Peter Tang · 2018
Later among the works it cites.
Fast approximate natural gradient descent in a kronecker factored eigenbasis
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas, and Pascal Vincent · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Kronecker-factored curvature approximations for recurrent neural networks
James Martens, Jimmy Ba, and Matt Johnson · 2018
Later among the works it cites.
Leonard Adolphs, Jonas Kohler, and Aurelien Lucchi · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Distributed second-order optimization using kronecker-factored approximations, 2016
Jimmy Ba, Roger Grosse, and James Martens · 2016
Cited alongside, same era.
A kronecker-factored approximate fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Cited alongside, same era.
Practical gauss-newton optimisation for deep learning
Aleksandar Botev, Hippolyt Ritter, and David Barber · 2017
Cited alongside, same era.
Sub-sampled cubic regularization for non-convex optimization
Jonas Moritz Kohler and Aurelien Lucchi · 2017
Cited alongside, same era.
Newton sketch: A near linear-time optimization algorithm with linear-quadratic convergence
Mert Pilanci and Martin J Wainwright · 2017
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Cited alongside, same era.
Limitations of the empirical fisher approximation for natural gradient descent
Frederik Kunstner, Philipp Hennig, and Lukas Balles · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Stochastic variance-reduced cubic regularization for nonconvex optimization
Zhe Wang, Yi Zhou, Yingbin Liang, and Guanghui Lan · 2019
Later among the works it cites.
Newton-type methods for non-convex optimization under inexact hessian information
Peng Xu, Fred Roosta, and Michael W Mahoney · 2019
Later among the works it cites.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model
Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George Dahl, Chris Shallue, and Roger B Grosse · 2019
Later among the works it cites.
Second-order optimization for non-convex machine learning: An empirical study
Peng Xu, Fred Roosta, and Michael W Mahoney · 2020
Closest in time.