Fetching the paper…
Reading the bibliography…
Gradient descent-based optimization methods underpin the parameter training of neural networks, and hence comprise a significant component in the impressive test results found in a number of applications.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris Polyak · 1964
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o(1/k2)
Yurii Nesterov · 1983
Earlier work this paper cites.
On the scope of the method of modified equations
DF Griffiths and JM Sanz-Serna · 1986
Earlier work this paper cites.
Parallel distributed processing: Explorations in the microstructure of cognition, vol. 1
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Stochastic relaxation, gibbs distributions, and the bayesian restoration of images
Stuart Geman and Donald Geman · 1987
Earlier work this paper cites.
Asymptotic global behavior for stochastic approximation and diffusions with slowly decreasing noise effects: global minimization via monte carlo
Harold J Kushner · 1987
Earlier work this paper cites.
Introduction to optimization
Boris T. Polyak · 1987
Earlier work this paper cites.
Experiments in nonconvex optimization: stochastic approximation with function smoothing and simulated annealing
MA Styblinski and T-S Tang · 1990
Earlier work this paper cites.
Simulated annealing
Dimitris Bertsimas, John Tsitsiklis, et al · 1993
Earlier work this paper cites.
Dynamical systems and numerical analysis , volume 2
Andrew Stuart and Anthony R Humphries · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey Hinton and Ruslan Salakhutdinov · 2006
Earlier work this paper cites.
Invariant manifolds , volume 583
Morris W Hirsch, Charles Chapman Pugh, and Michael Shub · 2006
Earlier work this paper cites.
Numerical integrators based on modified differential equations
Philippe Chartier, Ernst Hairer, and Gilles Vilmart · 2007
Earlier work this paper cites.
Multiscale Methods: Averaging and Homogenization , volume 53
Grigorios Pavliotis and Andrew Stuart · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Cited alongside, same era.
Convergence of numerical time-averaging and stationary measures via poisson equations
Jonathan C Mattingly, Andrew M Stuart, and Michael V Tretyakov · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Applications of centre manifold theory , volume 35
Jack Carr · 2012
Cited alongside, same era.
Stochastic approximation methods for constrained and unconstrained systems , volume 26
Harold Joseph Kushner and Dean S Clark · 2012
Cited alongside, same era.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
A lyapunov analysis of momentum methods in optimization
Ashia C. Wilson, Benjamin Recht, and Michael I. Jordan · 2016
Later among the works it cites.
Unified convergence analysis of stochastic momentum methods for convex and non-convex optimization
Tianbao Yang, Qihang Lin, and Zhe Li · 2016
Later among the works it cites.
Bridging the gap between constant step size stochastic gradient descent and markov chains
Aymeric Dieuleveut, Alain Durmus, and Francis Bach · 2017
Later among the works it cites.
Dissipativity theory for nesterov’s accelerated method
Bin Hu and Laurent Lessard · 2017
Later among the works it cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and Weinan E · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
Normally hyperbolic invariant manifolds in dynamical systems , volume 105
Stephen Wiggins · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Introductory Lectures on Convex Optimization: A Basic Course
Yurii Nesterov · 2014
Cited alongside, same era.
A differential equation for modeling nesterov’s accelerated gradient method: Theory and insights
Weijie Su, Stephen Boyd, and Emmanuel Candes · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Cited alongside, same era.
Linearly convergent stochastic heavy ball method for minimizing generalization error
Nicolas Loizou and Peter Richtárik · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
Integration methods and optimization algorithms
Damien Scieur, Vincent Roulet, Francis Bach, and Alexandre d’Aspremont · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
On symplectic optimization, 2018
Michael Betancourt, Michael I. Jordan, and Ashia C. Wilson · 2018
Later among the works it cites.
Multiscale analysis of accelerated gradient methods
Mohammad Farazmand · 2018
Later among the works it cites.
Semigroups of stochastic gradient descent and online principal component analysis: properties and diffusion approximations
Yuanyuan Feng, Lei Li, and Jian-Guo Liu · 2018
Later among the works it cites.
Stochastic heavy ball
Sébastien Gadat, Fabien Panloup, Sofiane Saadane, et al · 2018
Later among the works it cites.
Understanding the acceleration phenomenon via high-resolution differential equations
Bin Shi, Simon S Du, Michael I Jordan, and Weijie J Su · 2018
Later among the works it cites.
Direct runge-kutta discretization achieves acceleration
Jingzhao Zhang, Aryan Mokhtari, Suvrit Sra, and Ali Jadbabaie · 2018
Later among the works it cites.
Multiscale analysis of accelerated gradient methods
Mohammad Farazmand · 2020
Closest in time.
Continuous time analysis of momentum methods
Nikola B. Kovachki and Andrew M. Stuart · 2021
Closest in time.