Fetching the paper…
Reading the bibliography…
While momentum-based accelerated variants of stochastic gradient descent (SGD) are widely used when training machine learning models, there is little theoretical understanding on the generalization error of such methods.
Some methods of speeding up the convergence of iteration methods
Boris T. Polyak · 1964
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
Vladimir N. Vapnik and Alexey Ya Chervonenkis · 1971
Earlier work this paper cites.
Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables
Milton Abramowitz and Irene A. Stegun · 1972
Earlier work this paper cites.
Monotone operators and the proximal point algorithm
R. Tyrrell Rockafellar · 1976
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
Arkadij Nemirovskij and David Borisovich Yudin · 1983
Earlier work this paper cites.
A method of solving a convex programming problem with convergence O ( 1 / k 2 ) O(1/k^{2})
Yurii Nesterov · 1983
Earlier work this paper cites.
A theory of the learnable
Leslie G. Valiant · 1984
Earlier work this paper cites.
Learning representations by backpropagating errors
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1986
Earlier work this paper cites.
Incorporating second-order functional knowledge for better option pricing
Charles Dugas, Yoshua Bengio, François Bélisle, Claude Nadeau, and René Garcia · 2000
Earlier work this paper cites.
Localized rademacher complexities
Peter L Bartlett, Olivier Bousquet, and Shahar Mendelson · 2002
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
New exponential bounds and approximations for the computation of error probability in fading channels
Marco Chiani, Davide Dardari, and Marvin K. Simon · 2003
Earlier work this paper cites.
On the complexity of linear prediction: risk bounds, margin bounds, and regularization
Sham M. Kakade, Karthik Sridharan, and Ambuj Tewari · 2008
Earlier work this paper cites.
Fast rates for regularized objectives
Karthik Sridharan, Shai Shalev-Shwartz, and Nathan Srebro · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Learnability, stability and uniform convergence
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan · 2010
Cited alongside, same era.
ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey E. Hinton · 2013
Cited alongside, same era.
Table of Integrals, Series, and Products
Izrail Solomonovich Gradshteyn and Iosif Moiseevich Ryzhik · 2014
Cited alongside, same era.
Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm
Deanna Needell, Nati Srebro, and Rachel Ward · 2014
Cited alongside, same era.
iPiano: Inertial proximal algorithm for nonconvex optimization
Peter Ochs, Yunjin Chen, Thomas Brox, and Thomas Pock · 2014
Analysis and design of optimization algorithms via integral quadratic constraints
Laurent Lessard, Benjamin Recht, and Andrew Packard · 2016
Later among the works it cites.
Understanding generalization
Ming Yang Ong · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C. Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
Generalization bounds for uniformly stable algorithms
Vitaly Feldman and Jan Vondrak · 2018
Closest in time.
Stochastic heavy ball
Sébastien Gadat, Fabien Panloup, and Sofiane Saadane · 2018
Closest in time.
A unified analysis of stochastic momentum methods for deep learning
Yan Yan, Tianbao Yang, Zhe Li, Qihang Lin, and Yi Yang · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Understanding Machine Learning: From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
A differential equation for modeling Nesterov’s acceleratedgradient method: Theory and insights
Weijie Su, Stephen Boyd, and Emmanuel Candès · 2014
Cited alongside, same era.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2015
Cited alongside, same era.
Global convergence of the heavy-ball method for convex optimization
Euhanna Ghadimi, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
iPiasco: Inertial proximal algorithm for strongly convex optimization
Peter Ochs, Thomas Brox, and Thomas Pock · 2015
Cited alongside, same era.
Closest in time.
Accelerated linear convergence of stochastic momentum methods in Wasserstein distances
Bugra Can, Mert Gürbüzbalaban, and Lingjiong Zhu · 2019
Closest in time.
High probability generalization bounds for uniformly stable algorithms with nearly optimal rate
Vitaly Feldman and Jan Vondrak · 2019
Closest in time.
On the convergence of Nesterov’s accelerated gradient method in stochastic settings
Mahmoud Assran and Michael Rabbat · 2020
Closest in time.
Sharper bounds for uniformly stable algorithms
Olivier Bousquet, Yegor Klochkov, and Nikita Zhivotovskiy · 2020
Closest in time.
A Lyapunov analysis for accelerated gradient methods: from deterministic to stochastic case
Maxime Laborde and Adam Oberman · 2020
Closest in time.
A high probability analysis of adaptive SGD with momentum
Xiaoyu Li and Francesco Orabona · 2020
Closest in time.
The role of memory in stochastic optimization
Antonio Orvieto, Jonas Kohler, and Aurelien Lucchi · 2020
Closest in time.
The instability of accelerated gradient descent
Amit Attia and Tomer Koren · 2021
Closest in time.
Stability and deviation optimal risk bounds with convergence rate O ( 1 / n ) O(1/n)
Yegor Klochkov and Nikita Zhivotovskiy · 2021
Closest in time.
A Lyapunov analysis of momentum methods in optimization
Ashia C. Wilson, Ben Recht, and Michael I. Jordan · 2021
Closest in time.