Fetching the paper…
Reading the bibliography…
Following the same routine as [SSJ20], we continue to present the theoretical analysis for stochastic gradient descent with momentum (SGD with momentum) in this paper.
Approximate integration of stochastic differential equations
G. N. Mil’shtein · 1975
Earlier work this paper cites.
Analyse numérique des équations différentielles stochastiques
D. Talay · 1982
Earlier work this paper cites.
Efficient numerical schemes for the approximation of expectations of functionals of the solution of a SDE and applications
D. Talay · 1984
Earlier work this paper cites.
Discretization and simulation of stochastic differential equations
E. Pardoux and D. Talay · 1985
Earlier work this paper cites.
Weak approximation of solutions of systems of stochastic differential equations
G. N. Mil’shtein · 1986
Earlier work this paper cites.
The approximation of multiple stochastic integrals
P. E. Kloeden and E. Platen · 1992
Earlier work this paper cites.
The law of the Euler scheme for stochastic differential equations: Ii. convergence rate of the density
V. Bally and D. Talay · 1996
Earlier work this paper cites.
On the trend to equilibrium for the Fokker-Planck equation: An interplay between physics and functional analysis
P. A. Markowich and C. Villani · 1999
Earlier work this paper cites.
Quantitative analysis of metastability in reversible diffusion processes via a Witten complex approach
B. Helffer, M. Klein, and F. Nier · 2004
Earlier work this paper cites.
Metastability in reversible diffusion processes II: Precise asymptotics for small eigenvalues
A. Bovier, V. Gayrard, and M. Klein · 2005
Earlier work this paper cites.
Hypoelliptic estimates and spectral theory for Fokker-Planck operators and Witten Laplacians
B. Helffer and F. Nier · 2005
Earlier work this paper cites.
Hypocoercive diffusion operators
C. Villani · 2006
Earlier work this paper cites.
Adaptive online gradient descent
E. Hazan, A. Rakhlin, and P. Bartlett · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A Krizhevsky · 2009
Cited alongside, same era.
Hypocoercivity
C Villani · 2009
Cited alongside, same era.
Tunnel effect and symmetries for Kramers–Fokker–Planck type operators
F. Hérau, M. Hitrik, and J. Sjöstrand · 2011
Cited alongside, same era.
Practical recommendations for gradient-based training of deep architectures
Y. Bengio · 2012
Cited alongside, same era.
An introduction to stochastic differential equations
L. Evans · 2012
Cited alongside, same era.
Global hypoellipticity and compactness of resolvent for fokker-planck operator
W.-X. Li · 2012
Cited alongside, same era.
A differential equation for modeling Nesterov’s accelerated gradient method: Theory and insights
W. Su, S. Boyd, and E. Candes · 2016
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Later among the works it cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Q. Li, C. Tai, and W. E · 2017
Later among the works it cites.
Cyclical learning rates for training neural networks
L. N. Smith · 2017
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Later among the works it cites.
Deep relaxation: partial differential equations for optimizing deep neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Stochastic Processes and Applications: Diffusion Processes, the Fokker–Planck and Langevin Equations
G. Pavliotis · 2014
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Gradient descent only converges to minimizers
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht · 2016
Cited alongside, same era.
A variational analysis of stochastic gradient algorithms
S. Mandt, M. Hoffman, and D. Blei · 2016
Cited alongside, same era.
P. Chaudhari, A. Oberman, S. Osher, S. Soatto, and G. Carlier · 2018
Later among the works it cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
P. Chaudhari and S. Soatto · 2018
Later among the works it cites.
Characterizing implicit bias in terms of optimization geometry
S. Gunasekar, J. Lee, D. Soudry, and N. Srebro · 2018
Later among the works it cites.
Dynamical, symplectic and stochastic perspectives on gradient-based optimization
M. I. Jordan · 2018
Later among the works it cites.
Understanding the acceleration phenomenon via high-resolution differential equations
B. Shi, S. Du, M. Jordan, and W. Su · 2018
Later among the works it cites.
Gradient flow algorithms for density propagation in stochastic systems
K. Caluya and A. Halder · 2019
Later among the works it cites.
On learning rates and schrödinger operators
B. Shi, W. Su, and M. Jordan · 2020
Later among the works it cites.