Fetching the paper…
Reading the bibliography…
We provide sharp path-dependent generalization and excess risk guarantees for the full-batch Gradient Descent (GD) algorithm on smooth losses (possibly non-Lipschitz, possibly nonconvex).
On generalization error bounds of noisy gradient methods for non-convex learning
Jian Li, Xuanyuan Luo, and Mingda Qiao · 1902
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Introductory lectures on convex programming, 1998
Yu Nesterov · 1998
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Convergence of the iterates of descent methods for analytic cost functions
P. A. Absil, R. Mahony, and B. Andrews · 2005
Earlier work this paper cites.
High probability convergence and uniform stability bounds for nonconvex stochastic gradient descent
Liam Madden, Emiliano Dall’Anese, and Stephen Becker · 2006
Earlier work this paper cites.
Smoothness, low noise and fast rates
Nathan Srebro, Karthik Sridharan, and Ambuj Tewari · 2010
Earlier work this paper cites.
Generalization of erm in stochastic convex optimization: The dimension strikes back
Vitaly Feldman · 2016
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the Polyak-Łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
Gradient descent only converges to minimizers
Jason D. Lee, Max Simchowitz, Michael I. Jordan, and Benjamin Recht · 2016
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Stability and generalization of learning algorithms that converge to global optima
Zachary Charles and Dimitris Papailiopoulos · 2018
Earlier work this paper cites.
Generalization bounds for uniformly stable algorithms
Vitaly Feldman and Jan Vondrak · 2018
Earlier work this paper cites.
Data-dependent stability of stochastic gradient descent
Ilja Kuzborskij and Christoph Lampert · 2018
Earlier work this paper cites.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Earlier work this paper cites.
Generalization bounds of SGLD for non-convex learning: Two theoretical viewpoints
Wenlong Mou, Liwei Wang, Xiyu Zhai, and Kai Zheng · 2018
Cited alongside, same era.
Generalization error bounds for noisy, iterative algorithms
Ankit Pensia, Varun Jog, and Po-Ling Loh · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
High probability generalization bounds for uniformly stable algorithms with nearly optimal rate
Vitaly Feldman and Jan Vondrak · 2019
Cited alongside, same era.
SGD: General analysis and improved rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik · 2019
Cited alongside, same era.
Information-theoretic generalization bounds for SGLD via data-dependent estimates
Sharper generalization bounds for learning with gradient-dominated objective functions
Yunwen Lei and Yiming Ying · 2021
Later among the works it cites.
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2021
Later among the works it cites.
Information-theoretic generalization bounds for stochastic gradient descent
Gergely Neu, Gintare Karolina Dziugaite, Mahdi Haghifam, and Daniel M. Roy · 2021
Later among the works it cites.
On the generalization of stochastic gradient descent with momentum
Ali Ramezani-Kebrya, Ashish Khisti, and Ben Liang · 2021
Later among the works it cites.
Optimizing information-theoretical generalization bound via anisotropic noise of SGLD
Bohan Wang, Huishuai Zhang, Jieyu Zhang, Qi Meng, Wei Chen, and Tie-Yan Liu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jeffrey Negrea, Mahdi Haghifam, Gintare Karolina Dziugaite, Ashish Khisti, and Daniel M. Roy · 2019
Cited alongside, same era.
Stability of stochastic gradient descent on nonsmooth convex losses
Raef Bassily, Vitaly Feldman, Cristóbal Guzmán, and Kunal Talwar · 2020
Cited alongside, same era.
Sharper generalization bounds for pairwise learning
Yunwen Lei, Antoine Ledent, and Marius Kloft · 2020
Cited alongside, same era.
Fine-grained analysis of stability and generalization for stochastic gradient descent
Yunwen Lei and Yiming Ying · 2020
Cited alongside, same era.
Never go full batch (in stochastic convex optimization)
Idan Amir, Yair Carmon, Tomer Koren, and Roi Livni · 2021
Cited alongside, same era.
Time-independent generalization bounds for SGLD in non-convex settings
Tyler Farghly and Patrick Rebeschini · 2021
Cited alongside, same era.
Stochastic training is not necessary for generalization
Jonas Geiping, Micah Goldblum, Phillip E Pope, Michael Moeller, and Tom Goldstein · 2021
Cited alongside, same era.
Analyzing the generalization capability of SGLD using properties of gaussian channels
Hao Wang, Yizhe Huang, Rui Gao, and Flavio Calmon · 2021
Later among the works it cites.
Stability and generalization for randomized coordinate descent
Puyu Wang, Liang Wu, and Yunwen Lei · 2021
Later among the works it cites.
On the algorithmic stability of adversarial training
Yue Xing, Qifan Song, and Guang Cheng · 2021
Later among the works it cites.
Stability of SGD: Tightness Analysis and Improved Bounds
Yikai Zhang, Wenjia Zhang, Sammy Bald, Vamsi Pingali, Chao Chen, and Mayank Goswami · 2021
Later among the works it cites.
Towards understanding why lookahead generalizes better than SGD and beyond
Pan Zhou, Hanshu Yan, Xiaotong Yuan, Jiashi Feng, and Shuicheng Yan · 2021
Later among the works it cites.
Understanding generalization error of SGD in nonconvex optimization
Yi Zhou, Yingbin Liang, and Huishuai Zhang · 2021
Later among the works it cites.
Convergence of gradient descent for deep neural networks
Sourav Chatterjee · 2022
Closest in time.
Bias reduction in sample-based optimization
Darinka Dentcheva and Yang Lin · 2022
Closest in time.
Generalization in supervised learning through Riemannian contraction
Leo Kozachkov, Patrick M Wensing, and Jean-Jacques Slotine · 2022
Closest in time.
Generalization bounds via convex analysis
Gergely Neu and Gábor Lugosi · 2022
Closest in time.
Understanding generalization error of SGD in nonconvex optimization
Yi Zhou, Yingbin Liang, and Huishuai Zhang · 2022
Closest in time.