Fetching the paper…
Reading the bibliography…
Momentum methods such as Polyak's heavy ball (HB) method, Nesterov's accelerated gradient (AG) as well as accelerated projected gradient (APG) method have been commonly used in machine learning practice, but their performance is quite sensitive to noise in the gradients.
The existence of stationary measures for certain Markov processes
T. E. Harris · 1956
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B.T. Polyak · 1964
Earlier work this paper cites.
Introduction to optimization
Boris T. Polyak · 1987
Earlier work this paper cites.
The n n th power of a 2 × 2 2\times 2 matrix
Kenneth S. Williams · 1992
Earlier work this paper cites.
Markov Chains and Stochastic Stability
S. P. Meyn and R. L. Tweedie · 1993
Earlier work this paper cites.
Computable bounds for geometric convergence rates of Markov chains
S. P. Meyn and R. L. Tweedie · 1994
Earlier work this paper cites.
Matrix computations
Gene H Golub and Charles F Van Loan · 1996
Earlier work this paper cites.
Optimization under uncertainty using momentum
Sjur Didrik Flåm · 2004
Earlier work this paper cites.
Introductory Lectures on Convex Optimization. Applied Optimization, Vol. 87
Yurii Nesterov · 2004
Earlier work this paper cites.
Signal recovery by proximal forward-backward splitting
Patrick L Combettes and Valérie R Wajs · 2005
Earlier work this paper cites.
Smooth optimization with approximate gradient
A. d’Aspremont · 2008
Earlier work this paper cites.
Information-theoretic lower bounds on the oracle complexity of convex optimization
Alekh Agarwal, Martin J Wainwright, Peter L. Bartlett, and Pradeep K. Ravikumar · 2009
Earlier work this paper cites.
Accelerated gradient methods for stochastic optimization and online learning
Chonghai Hu, Weike Pan, and James T Kwok · 2009
Earlier work this paper cites.
Matrix Iterative Analysis
Richard S Varga · 2009
Earlier work this paper cites.
Optimal Transport: Old and New
Cédric Villani · 2009
Earlier work this paper cites.
Dual averaging methods for regularized stochastic learning and online optimization
Lin Xiao · 2010
Earlier work this paper cites.
Probability and Stochastics
Erhan Çınlar · 2011
Earlier work this paper cites.
Yet another look at Harris’ ergodic theorem for Markov chains
M. Hairer and J. C. Mattingly · 2011
Earlier work this paper cites.
Information-based complexity, feedback and dynamics in convex programming
Maxim Raginsky and Alexander Rakhlin · 2011
Earlier work this paper cites.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization i: A generic algorithmic framework
Saeed Ghadimi and Guanghui Lan · 2012
Earlier work this paper cites.
An optimal method for stochastic composite optimization
Guanghui Lan · 2012
Cited alongside, same era.
Lyapunov analysis and the heavy ball method
Benjamin Recht · 2012
Cited alongside, same era.
Measurements-based power control-a cross-layered framework
Berk Birand, Howard Wang, Keren Bergman, and Gil Zussman · 2013
Cited alongside, same era.
Intermediate gradient methods for smooth convex problems with inexact oracle
O. Devolder, F. Glineur, and Y. Nesterov · 2013
Cited alongside, same era.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization, ii: Shrinking procedures and optimal algorithms
S. Ghadimi and G. Lan · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Analysis and design of optimization algorithms via integral quadratic constraints
Laurent Lessard, Benjamin Recht, and Andrew Packard · 2016
Later among the works it cites.
A Lyapunov analysis of momentum methods in optimization
A.C. Wilson, B. Recht, and M.I. Jordan · 2016
Later among the works it cites.
Unified convergence analysis of stochastic momentum methods for convex and non-convex optimization
Tianbao Yang, Qihang Lin, and Zhe Li · 2016
Later among the works it cites.
First-Order Methods in Optimization
A. Beck · 2017
Later among the works it cites.
Bridging the gap between constant step size stochastic gradient descent and Markov chains
Aymeric Dieuleveut, Alain Durmus, and Francis Bach · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The Nature of Statistical Learning Theory
Vladimir Vapnik · 2013
Cited alongside, same era.
Private empirical risk minimization: Efficient algorithms and tight error bounds
Raef Bassily, Adam Smith, and Abhradeep Thakurta · 2014
Cited alongside, same era.
Theory of Convex Optimization for Machine Learning
S. Bubeck · 2014
Cited alongside, same era.
First-order methods of smooth convex optimization with inexact oracle
O. Devolder, F. Glineur, and Y. Nesterov · 2014
Cited alongside, same era.
Global convergence of the Heavy-ball method for convex optimization
Euhanna Ghadimi, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2014
Cited alongside, same era.
Robustness versus acceleration., August 2014
M. Hardt · 2014
Cited alongside, same era.
Harder, better, faster, stronger convergence rates for least-squares regression
Aymeric Dieuleveut, Nicolas Flammarion, and Francis Bach · 2017
Later among the works it cites.
A dynamical systems perspective to convergence rate analysis of proximal algorithms
Mahyar Fazlyab, Alejandro Ribeiro, Manfred Morari, and Victor M Preciado · 2017
Later among the works it cites.
Dissipativity theory for Nesterov’s accelerated method
B. Hu and L. Lessard · 2017
Later among the works it cites.
Accelerating stochastic gradient descent
Prateek Jain, Sham M Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford · 2017
Later among the works it cites.
Sahar Karimi and Stephen Vavasis · 2017
Later among the works it cites.
Nicolas Loizou and Peter Richtárik · 2017
Later among the works it cites.
Non-convex learning via stochastic gradient Langevin dynamics: a nonasymptotic analysis
M. Raginsky, A. Rakhlin, and M. Telgarsky · 2017
Later among the works it cites.
Robust accelerated gradient methods for smooth strongly convex functions
N. S. Aybat, A. Fallah, M. Gürbüzbalaban, and A. Ozdaglar · 2018
Later among the works it cites.
On Acceleration with Noise-Corrupted Gradients
Michael B. Cohen, Jelena Diakonikolas, and Lorenzo Orecchia · 2018
Later among the works it cites.
Breaking Reversibility Accelerates Langevin Dynamics for Global Non-Convex Optimization
Xuefeng Gao, Mert Gurbuzbalaban, and Lingjiong Zhu · 2018
Later among the works it cites.
Xuefeng Gao, Mert Gürbüzbalaban, and Lingjiong Zhu · 2018
Later among the works it cites.
Stochastic heavy ball
Sébastien Gadat, Fabien Panloup, and Sofiane Saadane · 2018
Later among the works it cites.
Accelerated gossip via stochastic heavy ball method
Nicolas Loizou and Peter Richtárik · 2018
Later among the works it cites.
A universally optimal multistage accelerated stochastic gradient method
Necdet Serhat Aybat, Alireza Fallah, Mert Gurbuzbalaban, and Asuman Ozdaglar · 2019
Closest in time.
A tail-index analysis of stochastic gradient noise in deep neural networks
Umut Simsekli, Levent Sagun, and Mert Gurbuzbalaban · 2019
Closest in time.