Fetching the paper…
Reading the bibliography…
Stochastic gradient descent is one of the most common iterative algorithms used in machine learning and its convergence analysis is a rich area of research.
On tail probabilities for martingales
David A Freedman · 1975
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
Arkadi Nemirovsky and David Yudin · 1983
Earlier work this paper cites.
Probability and Measure
Patrick Billingsley · 1995
Earlier work this paper cites.
Gradient convergence in gradient methods with errors
D. Bertsekas and J. Tsitsiklis · 2000
Earlier work this paper cites.
The tradeoffs of large scale learning
Léon Bottou and Olivier Bousquet · 2008
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Exponential inequalities for martingales with applications
Xiequan Fan, Ion Grama, Quansheng Liu, et al · 2015
Earlier work this paper cites.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Saeed Ghadimi, Guanghui Lan, and Hongchao Zhang · 2016
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the Polyak-Łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
Proximal stochastic methods for nonsmooth nonconvex finite-sum optimization
Sashank J Reddi, Suvrit Sra, Barnabas Poczos, and Alexander J Smola · 2016
Earlier work this paper cites.
Guaranteed matrix completion via non-convex factorization
Ruoyu Sun and Zhi-Quan Luo · 2016
Earlier work this paper cites.
First-Order Methods in Optimization
A. Beck · 2017
Earlier work this paper cites.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Cited alongside, same era.
Stability and generalization of learning algorithms that converge to global optima
Zachary Charles and Dimitris Papailiopoulos · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Cited alongside, same era.
Lectures on Convex Optimization , volume 137 of Springer Optimization and Its Applications
Y. Nesterov · 2018
Cited alongside, same era.
High-Dimensional Probability: An Introduction with Applications in Data Science , volume 47
Roman Vershynin · 2018
Cited alongside, same era.
Robustness analysis of non-convex stochastic gradient descent using biased expectations
Kevin Scaman and Cedric Malherbe · 2020
Closest in time.
Sub-Weibull distributions: Generalizing sub-Gaussian and sub-exponential properties to heavier tailed distributions
Mariia Vladimirova, Stéphane Girard, Hien Nguyen, and Julyan Arbel · 2020
Closest in time.
Lasso guarantees for β \beta -mixing heavy-tailed time series
Kam Chung Wong, Zifan Li, and Ambuj Tewari · 2020
Closest in time.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie · 2020
Closest in time.
High-probability bounds for non-convex stochastic optimization with heavy tails
Ashok Cutkosky and Harsh Mehta · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yossi Arjevani, Yair Carmon, John C Duchi, Dylan J Foster, Nathan Srebro, and Blake Woodworth · 2019
Cited alongside, same era.
A tail-index analysis of stochastic gradient noise in deep neural networks
Umut Şimşekli, Levent Sagun, and Mert Gürbüzbalaban · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon S Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
Tight analyses for non-smooth stochastic gradient descent
Nicholas J. A. Harvey, Christopher Liaw, Yaniv Plan, and Sikander Randhawa · 2019
Cited alongside, same era.
Continuous-time models for stochastic optimization algorithms
Antonio Orvieto and Aurelien Lucchi · 2019
Cited alongside, same era.
Non-gaussianity of stochastic gradient noise
Abhishek Panigrahi, Raghav Somani, Navin Goyal, and Praneeth Netrapalli · 2019
Cited alongside, same era.
Better theory for SGD in the nonconvex world
Ahmed Khaled and Peter Richtárik · 2020
Cited alongside, same era.
Mert Gürbüzbalaban, Umut Şimşekli, and Lingjiong Zhu · 2021
Closest in time.
Stopping criteria for, and strong convergence of, stochastic gradient descent on Bottou-Curtis-Nocedal functions
Vivak Patel · 2021
Closest in time.
Almost sure convergence rates for stochastic gradient descent and stochastic heavy ball
Othmane Sebbouh, Robert M Gower, and Aaron Defazio · 2021
Closest in time.
High probability guarantees for nonconvex stochastic gradient descent with heavy tails
Shaojie Li and Yong Liu · 2022
Closest in time.
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2022
Closest in time.
Global convergence and stability of stochastic gradient descent
Vivak Patel, Shushu Zhang, and Bowen Tian · 2022
Closest in time.
Sharp concentration results for heavy-tailed distributions
Milad Bakhshizadeh, Arian Maleki, and Victor H de la Pena · 2023
Closest in time.
Memory capacity of two layer neural networks with smooth activation
Liam Madden and Christos Thrampoulidis · 2024
Closest in time.