Fetching the paper…
Reading the bibliography…
We present a generalization of Nesterov's accelerated gradient descent algorithm.
Unified Optimal Analysis of the (Stochastic) Gradient Method
Sebastian U. Stich · 1907
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Methods of conjugate gradients for solving linear systems
Magnus R Hestenes and Eduard Stiefel · 1952
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Yurii Nesterov · 1983
Earlier work this paper cites.
The variational formulation of the fokker–planck equation
Richard Jordan, David Kinderlehrer, and Felix Otto · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
A second-order gradient-like dissipative dynamical system with Hessian-driven damping: Application to optimization and mechanics
Felipe Alvarez, Hedy Attouch, Jérôme Bolte, and Patrick Redont · 2002
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Amir Beck and Marc Teboulle · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization i: A generic algorithmic framework
Saeed Ghadimi and Guanghui Lan · 2012
Earlier work this paper cites.
Efficiency of coordinate descent methods on huge-scale optimization problems
Yurii Nesterov · 2012
Earlier work this paper cites.
Probability Theory: A Comprehensive Course
A. Klenke · 2013
Earlier work this paper cites.
Gradient methods for minimizing composite functions
Yurii Nesterov · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Variational modelling: Energies, gradient flows, and large deviations
Mark A Peletier · 2014
Earlier work this paper cites.
A differential equation for modeling nesterov’s accelerated gradient method: theory and insights
Weijie Su, Stephen Boyd, and Emmanuel Candes · 2014
Earlier work this paper cites.
On the convergence of the iterates of FISTA
Antonin Chambolle and Charles H Dossal · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the Polyak-Lojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
A variational perspective on accelerated methods in optimization
Andre Wibisono, Ashia C Wilson, and Michael I Jordan · 2016
Cited alongside, same era.
On exponential convergence of SGD in non-convex over-parametrized learning
Raef Bassily, Mikhail Belkin, and Siyuan Ma · 2018
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaïc Chizat and Francis Bach · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Cited alongside, same era.
Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron
Sharan Vaswani, Francis Bach, and Mark Schmidt · 2019
Later among the works it cites.
Global convergence of adaptive gradient methods for an over-parameterized neural network
Xiaoxia Wu, Simon S Du, and Rachel Ward · 2019
Later among the works it cites.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2020
Later among the works it cites.
A simpler approach to accelerated optimization: iterative averaging meets optimism
Pooria Joulani, Anant Raj, Andras Gyorgy, and Csaba Szepesvári · 2020
Later among the works it cites.
A lyapunov analysis for accelerated gradient methods: from deterministic to stochastic case
Maxime Laborde and Adam Oberman · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Universal method for stochastic composite optimization problems
Alexander Vladimirovich Gasnikov and Yu E Nesterov · 2018
Cited alongside, same era.
Accelerating stochastic gradient descent for least squares regression
Prateek Jain, Sham M Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford · 2018
Cited alongside, same era.
On the insufficiency of existing momentum schemes for stochastic optimization
Rahul Kidambi, Praneeth Netrapalli, Prateek Jain, and Sham Kakade · 2018
Cited alongside, same era.
Another look at the fast iterative shrinkage/thresholding algorithm (fista)
Donghwan Kim and Jeffrey A Fessler · 2018
Cited alongside, same era.
Online adaptive methods, universality and acceleration
Kfir Y Levy, Alp Yurtsever, and Volkan Cevher · 2018
Cited alongside, same era.
Accelerating sgd with momentum for over-parameterized learning
Chaoyue Liu and Mikhail Belkin · 2018
Cited alongside, same era.
Fixing weight decay regularization in adam, 2018
Ilya Loshchilov and Frank Hutter · 2018
Cited alongside, same era.
Towards theoretically understanding why sgd generalizes better than adam in deep learning
Pan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong, Steven Hoi, and E. Weinan · 2020
Later among the works it cites.
Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation
Mikhail Belkin · 2021
Later among the works it cites.
A continuized view on nesterov acceleration
Raphaël Berthier, Francis Bach, Nicolas Flammarion, Pierre Gaillard, and Adrien Taylor · 2021
Later among the works it cites.
Label noise SGD provably prefers flat global minimizers
Alex Damian, Tengyu Ma, and Jason D. Lee · 2021
Later among the works it cites.
Generalized momentum-based methods: A hamiltonian perspective
Jelena Diakonikolas and Michael I Jordan · 2021
Later among the works it cites.
A continuized view on Nesterov acceleration for stochastic gradient descent and randomized gossip
Mathieu Even, Raphaël Berthier, Francis Bach, Nicolas Flammarion, Pierre Gaillard, Hadrien Hendrikx, Laurent Massoulié, and Adrien Taylor · 2021
Later among the works it cites.
Understanding the acceleration phenomenon via high-resolution differential equations
Bin Shi, Simon S Du, Michael I Jordan, and Weijie J Su · 2021
Later among the works it cites.
A lyapunov analysis of accelerated methods in optimization
Ashia C Wilson, Ben Recht, and Michael I Jordan · 2021
Later among the works it cites.
Stochastic gradient descent with noise of machine learning type. Part II: Continuous time analysis
Stephan Wojtowytsch · 2021
Later among the works it cites.
First-order optimization algorithms via inertial systems with hessian driven damping
Hedy Attouch, Zaki Chbani, Jalal Fadili, and Hassan Riahi · 2022
Later among the works it cites.
On the fast convergence of minibatch heavy ball momentum
Raghu Bollapragada, Tyler Chen, and Rachel Ward · 2022
Later among the works it cites.
Stochastic differential equations for modeling first order optimization methods
Marc Dambrine, Ch Dossal, Bénédicte Puig, and Aude Rondepierre · 2022
Later among the works it cites.
What happens after SGD reaches zero loss? –a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2022
Later among the works it cites.
The error-feedback framework: Better rates for sgd with delayed gradients and compressed updates
Sebastian U. Stich and Sai Praneeth Karimireddy · 2022
Later among the works it cites.
Continuous-time analysis of accelerated gradient methods via conservation laws in dilated coordinate systems
Jaewook J Suh, Gyumin Roh, and Ernest K Ryu · 2022
Later among the works it cites.
Jonathan W Siegel and Stephan Wojtowytsch · 2023
Closest in time.
Stochastic gradient descent with noise of machine learning type. Part I: Discrete time analysis
Stephan Wojtowytsch · 2023
Closest in time.
Train cifar10 with pytorch
Kuang Liu · 2024
Closest in time.