Fetching the paper…
Reading the bibliography…
First-order methods such as stochastic gradient descent (SGD) are currently the standard algorithm for training deep neural networks.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
Fixup initialization: Residual learning without normalization
Hongyi Zhang, Yann N Dauphin, and Tengyu Ma · 1901
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 1904
Earlier work this paper cites.
Fast convergence of natural gradient descent for overparameterized neural networks
Guodong Zhang, James Martens, and Roger Grosse · 1905
Earlier work this paper cites.
A method for the solution of certain non-linear problems in least squares
Kenneth Levenberg · 1944
Earlier work this paper cites.
Numerical methods for solving linear least squares problems
Gene Golub · 1965
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
Sue Becker, Yann Le Cun, et al · 1988
Earlier work this paper cites.
Matrix computations (3rd ed.)
Gene H. Golub and Charles F. Van Loan · 1996
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
The method of subspace corrections
Jinchao Xu · 2001
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Escaping from saddle points − - online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
A kronecker-factored approximate fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Ordinal regression with multiple output cnn for age estimation
Zhenxing Niu, Mo Zhou, Le Wang, Xinbo Gao, and Gang Hua · 2016
Cited alongside, same era.
http://rsnachallenges.cloudapp.net/competitions/4, 2017
Revisiting small batch training for deep neural networks
Dominic Masters and Carlo Luschi · 2018
Later among the works it cites.
Foundations of machine learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2018
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2018
Later among the works it cites.
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rsna pediatric bone age challenge · 2017
Cited alongside, same era.
Practical gauss-newton optimisation for deep learning
Aleksandar Botev, Hippolyt Ritter, and David Barber · 2017
Cited alongside, same era.
Generalization bounds of sgld for non-convex learning: Two theoretical viewpoints
Wenlong Mou, Liwei Wang, Xiyu Zhai, and Kai Zheng · 2017
Cited alongside, same era.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Yuhuai Wu, Elman Mansimov, Roger B Grosse, Shun Liao, and Jimmy Ba · 2017
Cited alongside, same era.
Exact natural gradient in deep linear networks and its application to the nonlinear case
Alberto Bernacchia, Máté Lengyel, and Guillaume Hennequin · 2018
Cited alongside, same era.
A note on lazy training in supervised differentiable programming
Lenaic Chizat and Francis Bach · 2018
Cited alongside, same era.
Fast approximate natural gradient descent in a kronecker factored eigenbasis
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas, and Pascal Vincent · 2018
Cited alongside, same era.
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Closest in time.
Xialiang Dou and Tengyuan Liang · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Closest in time.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2019
Closest in time.
Adaptive gradient methods with dynamic bound of learning rate
Liangchen Luo, Yuanhao Xiong, Yan Liu, and Xu Sun · 2019
Closest in time.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Closest in time.
Efficient subsampled gauss-newton and natural gradient methods for training neural networks
Yi Ren and Donald Goldfarb · 2019
Closest in time.
High-Dimensional Statistics: A Non-Asymptotic Viewpoint
Martin J. Wainwright · 2019
Closest in time.
Greg Yang · 2019
Closest in time.
An improved analysis of training over-parameterized deep neural networks
Difan Zou and Quanquan Gu · 2019
Closest in time.