Fetching the paper…
Reading the bibliography…
This is a handbook of simple proofs of the convergence of gradient and stochastic gradient descent type methods.
“Unified Optimal Analysis of the (Stochastic) Gradient Method”
Sebastian. Stich · 1907
Earlier work this paper cites.
“A Modern Introduction to Online Learning”
Francesco Orabona · 1912
Earlier work this paper cites.
“A stochastic approximation method”
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
“Some Methods of Speeding up the Convergence of Iteration Methods”
B.. Polyak · 1964
Earlier work this paper cites.
“Stochastic Optimization Problems with Nondifferentiable Cost Functionals”
Dimitri. Bertsekas · 1973
Earlier work this paper cites.
“On Cezari’s convergence of the steepest descent method for approximating saddle point of convex-concave functions”
Arkadi Nemirovski and David Yudin · 1978
Earlier work this paper cites.
“Problem complexity and method efficiency in optimization”
Arkadi Nemirovski and David. Yudin · 1983
Earlier work this paper cites.
“Generalized Hessian Matrix and Second-Order Optimality Conditions for Problems withC1,1 Data”
Jean-Baptiste Hiriart-Urruty, Jean-Jacques Strodiot and V. Nguyen · 1984
Earlier work this paper cites.
“Introduction to Optimization. Translations series in mathematics and engineering”
B.T. Polyak · 1987
Earlier work this paper cites.
“Optimization and Nonsmooth Analysis”, Classics in Applied Mathematics
F. Clarke · 1990
Earlier work this paper cites.
“Better Theory for SGD in the Nonconvex World”
Ahmed Khaled and Peter Richt“’arik · 2002
Earlier work this paper cites.
Simon Lacoste-Julien, Mark Schmidt and Francis Bach · 2002
Earlier work this paper cites.
“Online Convex Programming and Generalized Infinitesimal Gradient Ascent”
Martin Zinkevich · 2003
Earlier work this paper cites.
“Introductory Lectures on Convex Optimization”
Yurii Nesterov · 2004
Earlier work this paper cites.
“Unified Analysis of Stochastic Gradient Methods for Composite Convex and Smooth Optimization”
Ahmed Khaled, Othmane Sebbouh, Nicolas Loizou, Robert. Gower and Peter Richt“’arik · 2006
Earlier work this paper cites.
“Clarke Subgradients of Stratifiable Functions”
J. Bolte, A. Daniilidis, A. Lewis and M. Shiota · 2007
Earlier work this paper cites.
“Pegasos: primal estimated subgradient solver for SVM”
Shai Shalev-Shwartz, Yoram Singer and Nathan Srebro · 2007
Earlier work this paper cites.
“A Fast Iterative Shrinkage-Thresholding Algorithm for Linear Inverse Problems”
A. Beck and M. Teboulle · 2009
Cited alongside, same era.
“Robust stochastic approximation approach to stochastic programming”
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan and Alexander Shapiro · 2009
Cited alongside, same era.
“Variational Analysis”
R. Rockafellar and Roger J.-B. Wets · 2009
Cited alongside, same era.
“Non-asymptotic analysis of stochastic approximation algorithms for machine learning”
Eric Moulines and Francis Bach · 2011
Cited alongside, same era.
“Convergence Rates of Inexact Proximal-Gradient Methods for Convex Optimization”
Mark Schmidt, Nicolas Roux and Francis Bach · 2011
Cited alongside, same era.
“Fast Convergence of Stochastic Gradient Descent under a Strong Growth Condition”
Mark Schmidt and Nicolas Roux · 2013
“The Power of Interpolation: Understanding the Effectiveness of SGD in Modern Over-parametrized Learning”
Siyuan Ma, Raef Bassily and Mikhail Belkin · 2018
Later among the works it cites.
“Stochastic (Approximate) Proximal Point Methods: Convergence, Optimality, and Adaptivity”
Hilal Asi and John. Duchi · 2019
Later among the works it cites.
“Stochastic model-based minimization of weakly convex functions”
Damek Davis and Dmitriy Drusvyatskiy · 2019
Later among the works it cites.
“Optimal mini-batch and step sizes for SAGA”
Nidham Gazagnadou, Robert Gower and Joseph Salmon · 2019
Later among the works it cites.
“SGD: General Analysis and Improved Rates”
Robert Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin and Peter Richt“’arik · 2019
Later among the works it cites.
“Training Neural Networks for and by Interpolation”
Leonard Berrada, Andrew Zisserman and M. Kumar · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Stochastic Gradient Descent for Non-smooth Optimization: Convergence Results and Optimal Averaging Schemes”
Ohad Shamir and Tong Zhang · 2013
Cited alongside, same era.
“A Course in Mathematical Analysis: Volume 2, Metric and Topological Spaces, Functions of a Vector Variable”
D… Garling · 2014
Cited alongside, same era.
“Convex Optimization: Algorithms and Complexity”
S“’ebastien Bubeck · 2015
Cited alongside, same era.
“Global convergence of the Heavy-ball method for convex optimization”
Euhanna Ghadimi, Hamid Feyzmahdavian and Mikael Johansson · 2015
Cited alongside, same era.
“Convex Optimization in Normed Spaces”, SpringerBriefs in Optimization
Juan Peypouquet · 2015
Cited alongside, same era.
“Sketch and Project: Randomized Iterative Methods for Linear Systems and Inverting Matrices”
Robert. Gower · 2016
Cited alongside, same era.
Later among the works it cites.
“A Unified Theory of SGD: Variance Reduction, Sampling, Quantization and Coordinate Descent”
Eduard Gorbunov, Filip Hanzely and Peter Richt“’arik · 2020
Later among the works it cites.
“Toward a theory of optimization for over-parameterized systems of non-linear equations: the lessons of deep learning”
Chaoyue Liu, Libin Zhu and Mikhail Belkin · 2020
Later among the works it cites.
“Factorial Powers for Stochastic Optimization”
Aaron Defazio and Robert. Gower · 2021
Later among the works it cites.
“Stochastic Quasi-Gradient Methods: Variance Reduction via Jacobian Sketching”
Robert. Gower, Peter Richt“’arik and Francis Bach · 2021
Later among the works it cites.
“SGD for Structured Nonconvex Functions: Learning Rates, Minibatching and Interpolation”
Robert. Gower, Othmane Sebbouh and Nicolas Loizou · 2021
Later among the works it cites.
“Stochastic Polyak Step-Size for Sgd: An Adaptive Learning Rate for Fast Convergence”
Nicolas Loizou, Sharan Vaswani, Issam Laradji and Simon Lacoste-Julien · 2021
Later among the works it cites.
“Almost sure convergence rates for Stochastic Gradient Descent and Stochastic Heavy Ball”
Othmane Sebbouh, Robert. Gower and Aaron Defazio · 2021
Later among the works it cites.
“Square Distance Functions Are Polyak-Łojasiewicz and Vice-Versa”
Guillaume Garrigos · 2023
Closest in time.
Guillaume Garrigos, Robert. Gower and Fabian Schaipp · 2023
Closest in time.
“Convergence of the Forward-Backward Algorithm: Beyond the Worst-Case with the Help of Geometry”
Guillaume Garrigos, Lorenzo Rosasco and Silvia Villa · 2023
Closest in time.
“Stochastic Polyak Step-size, a simple step-size tuner with optimal rates”, http://fa.bianp.net/blog/2023/sps/ , 2023
Fabian Pedregosa · 2023
Closest in time.