Fetching the paper…
Reading the bibliography…
This paper studies an intriguing phenomenon related to the good generalization performance of estimators obtained by using large learning rates within gradient descent algorithms.
Theory of reproducing kernels
Nachman Aronszajn · 1950
Earlier work this paper cites.
Classification under polynomial entropy and margin assumptions and randomized estimators
J-Y Audibert · 2004
Earlier work this paper cites.
Learning from examples as an inverse problem
Ernesto De Vito, Lorenzo Rosasco, Andrea Caponnetto, Umberto De Giovannini, and Francesca Odone · 2005
Earlier work this paper cites.
Fast learning rates for plug-in classifiers
Jean-Yves Audibert and Alexandre B Tsybakov · 2007
Earlier work this paper cites.
On regularization algorithms in learning theory
Frank Bauer, Sergei Pereverzev, and Lorenzo Rosasco · 2007
Earlier work this paper cites.
The tradeoffs of large scale learning
Léon Bottou and Olivier Bousquet · 2007
Earlier work this paper cites.
On early stopping in gradient descent learning
Yuan Yao, Lorenzo Rosasco, and Andrea Caponnetto · 2007
Earlier work this paper cites.
Spectral algorithms for supervised learning
L. Lo Gerfo, L. Rosasco, F. Odone, E. De Vito, and A. Verri · 2008
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Self-concordant analysis for logistic regression
F. Bach · 2010
Cited alongside, same era.
Optimal rates for regularization of statistical inverse learning problems
Gilles Blanchard and Nicole Mücke · 2016
Cited alongside, same era.
End-to-End Kernel Learning with Supervised Convolutional Kernel Networks
Julien Mairal · 2016
Cited alongside, same era.
Deep convolutional neural networks for image classification: A comprehensive review
Waseem Rawat and Zenghui Wang · 2017
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
Lectures on Convex Optimization
Yurii Nesterov · 2018
Cited alongside, same era.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Later among the works it cites.
The break-even point on optimization trajectories of deep neural networks
Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, and Krzysztof Geras · 2020
Later among the works it cites.
Learning rate annealing can provably help generalization, even for convex problems
Preetum Nakkiran · 2020
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
Later among the works it cites.
Stochastic training is not necessary for generalization
Jonas Geiping, Micah Goldblum, Phillip E. Pope, Michael Moeller, and Tom Goldstein · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen Xing, Devansh Arpit, Christos Tsirigotis, and Yoshua Bengio · 2018
Cited alongside, same era.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Statistical Optimality of Stochastic Gradient Descent on Hard Learning Problems through Multiple Passes
Loucas Pillaud-Vivien, Alessandro Rudi, and Francis Bach
Cited in the paper.
Exponential convergence of testing error for stochastic gradient methods
Loucas Pillaud-Vivien, Alessandro Rudi, and Francis Bach
Cited in the paper.
Later among the works it cites.
Catastrophic Fisher explosion: Early phase Fisher matrix impacts generalization
Stanislaw Jastrzebski, Devansh Arpit, Oliver Astrand, Giancarlo B Kerg, Huan Wang, Caiming Xiong, Richard Socher, Kyunghyun Cho, and Krzysztof J Geras · 2021
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
Samuel L Smith, Benoit Dherin, David Barrett, and Soham De · 2021
Later among the works it cites.
Direction matters: On the implicit bias of stochastic gradient descent with moderate learning rate
Jingfeng Wu, Difan Zou, Vladimir Braverman, and Quanquan Gu · 2021
Later among the works it cites.