Fetching the paper…
Reading the bibliography…
We consider the development of practical stochastic quasi-Newton, and in particular Kronecker-factored block-diagonal BFGS and L-BFGS methods, for training deep neural networks (DNNs).
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
The convergence of a class of double-rank minimization algorithms 1. general considerations
C. G. Broyden · 1970
Earlier work this paper cites.
A new approach to variable metric algorithms
R. Fletcher · 1970
Earlier work this paper cites.
A family of variable-metric methods derived by variational means
D. Goldfarb · 1970
Earlier work this paper cites.
Conditioning of quasi-newton methods for function minimization
D. F. Shanno · 1970
Earlier work this paper cites.
Algorithms for nonlinear constraints that use lagrangian functions
M. J. Powell · 1978
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
D. C. Liu and J. Nocedal · 1989
Earlier work this paper cites.
Representations of quasi-newton matrices and their use in limited memory methods
R. H. Byrd, J. Nocedal, and R. B. Schnabel · 1994
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Adaptive method of realizing natural gradient learning for multilayer perceptrons
S.-I. Amari, H. Park, and K. Fukumizu · 2000
Earlier work this paper cites.
On "natural" learning and pruning in multilayered perceptrons
T. Heskes · 2000
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
G. E. Hinton and R. R. Salakhutdinov · 2006
Earlier work this paper cites.
Numerical optimization
J. Nocedal and S. Wright · 2006
Earlier work this paper cites.
A stochastic quasi-newton method for online convex optimization
N. N. Schraudolph, J. Yu, and S. Günter · 2007
Earlier work this paper cites.
Sgd-qn: Careful quasi-newton stochastic gradient descent
A. Bordes, L. Bottou, and P. Gallinari · 2009
Earlier work this paper cites.
Deep learning via hessian-free optimization
J. Martens · 2010
Cited alongside, same era.
A fast natural newton method
N. L. Roux and A. W. Fitzgibbon · 2010
Cited alongside, same era.
On the use of stochastic hessian information in optimization methods for machine learning
R. H. Byrd, G. M. Chin, W. Neveitt, and J. Nocedal · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
G. Hinton, N. Srivastava, and K. Swersky · 2012
Cited alongside, same era.
Krylov subspace descent for deep learning
O. Vinyals and D. Povey · 2012
Cited alongside, same era.
A stochastic quasi-newton method for large-scale optimization
R. H. Byrd, S. L. Hansen, J. Nocedal, and Y. Singer · 2016
Later among the works it cites.
Stochastic block bfgs: Squeezing more curvature out of data
R. Gower, D. Goldfarb, and P. Richtárik · 2016
Later among the works it cites.
A linearly-convergent stochastic l-bfgs algorithm
P. Moritz, R. Nishihara, and M. Jordan · 2016
Later among the works it cites.
Distributed second-order optimization using kronecker-factored approximations
J. Ba, R. B. Grosse, and J. Martens · 2017
Later among the works it cites.
Practical gauss-newton optimisation for deep learning
A. Botev, H. Ritter, and D. Barber · 2017
Later among the works it cites.
Randomized quasi-newton updates are linearly convergent matrix inversion algorithms
R. M. Gower and P. Richtárik · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
Sequential quadratic programming (sqp) for optimal control in direct numerical simulation of turbulent flow
H. Badreddine, S. Vandewalle, and J. Meyers · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Res: Regularized stochastic bfgs algorithm
A. Mokhtari and A. Ribeiro · 2014
Cited alongside, same era.
Natural neural networks
G. Desjardins, K. Simonyan, R. Pascanu, et al · 2015
Cited alongside, same era.
A variance reduced stochastic newton method
A. Lucchi, B. McWilliams, and T. Hofmann · 2015
Cited alongside, same era.
Stochastic quasi-newton methods for nonconvex stochastic optimization
X. Wang, S. Ma, D. Goldfarb, and W. Liu · 2017
Later among the works it cites.
A neural network model with bidirectional whitening
Y. Fujimoto and T. Ohira · 2018
Later among the works it cites.
Fast approximate natural gradient descent in a kronecker factored eigenbasis
T. George, C. Laurent, X. Bouthillier, N. Ballas, and P. Vincent · 2018
Later among the works it cites.
Shampoo: Preconditioned stochastic tensor optimization
V. Gupta, T. Koren, and Y. Singer · 2018
Later among the works it cites.
Natural gradient via optimal transport
W. Li and G. Montúfar · 2018
Later among the works it cites.
Modular block-diagonal curvature approximations for feedforward architectures
F. Dangel, P. Hennig, and S. Harmeling · 2019
Later among the works it cites.
Newton-type methods for non-convex optimization under inexact hessian information
P. Xu, F. Roosta, and M. W. Mahoney · 2019
Later among the works it cites.
Information newton’s flow: second-order optimization method in probability space
Y. Wang and W. Li · 2020
Closest in time.