Fetching the paper…
Reading the bibliography…
Using second-order optimization methods for training deep neural networks (DNNs) has attracted many researchers.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Quasi-Newton methods, motivation and theory
J. E. Dennis and Jorge J. Moré · 1977
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate O ( 1 / k 2 ) {O}(1/k^{2})
Yurii Nesterov · 1983
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Deep learning via Hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Hazan Elad, and Singer Yoram · 2011
Earlier work this paper cites.
On optimization methods for deep learning
Quoc V Le, Jiquan Ngiam, Adam Coates, Abhik Lahiri, Bobby Prochnow, and Andrew Y Ng · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Cited alongside, same era.
Training neural networks with stochastic Hessian-free optimization
Ryan Kiros · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
A kronecker-factored approximate fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Fast approximate natural gradient descent in a kronecker factored eigenbasis
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas, and Pascal Vincent · 2018
Later among the works it cites.
Kronecker-factored curvature approximations for recurrent neural networks
James Martens, Jimmy Ba, and Matt Johnson · 2018
Later among the works it cites.
Noisy natural gradient as variational inference
Guodong Zhang, Shengyang Sun, David Duvenaud, and Roger Grosse · 2018
Later among the works it cites.
Quasi-Newton methods for deep learning: Forget the past, just sample
Albert S Berahas, Majid Jahani, and Martin Takáč · 2019
Later among the works it cites.
Large-scale distributed second-order optimization using kronecker-factored approximate curvature for deep convolutional neural networks
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Distributed second-order optimization using kronecker-factored approximations
Jimmy Ba, Roger Grosse, and James Martens · 2017
Cited alongside, same era.
Accelerating Hessian-free Gauss-Newton full-waveform inversion via l-BFGS preconditioned conjugate-gradient algorithm
Wenyong Pan, Kristopher A Innanen, and Wenyuan Liao · 2017
Cited alongside, same era.
Eigenvalue corrected noisy natural gradient
Juhan Bae, Guodong Zhang, and Roger Grosse · 2018
Cited alongside, same era.
Three mechanisms of weight decay regularization
Guodong Zhang, Chaoqi Wang, Bowen Xu, and Roger Grosse · 2019
Later among the works it cites.
Kaixin Gao, Xiaolei Liu, Zhenghai Huang, Min Wang, Zidong Wang, Dachuan Xu, and Fan Yu · 2020
Closest in time.
Practical quasi-Newton methods for training deep neural networks
Donald Goldfarb, Yi Ren, and Achraf Bahamou · 2020
Closest in time.
Convolutional neural network training with distributed K-FAC
J. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu, and Ian T. Foster · 2020
Closest in time.
Sketchy empirical natural gradient methods for deep learning
Minghan Yang, Dong Xu, Yongfeng Li, Zaiwen Wen, and Mengyun Chen · 2020
Closest in time.