Fetching the paper…
Reading the bibliography…
The learning rate is perhaps the single most important parameter in the training of neural networks and, more broadly, in stochastic (nonconvex) optimization.
Approximate integration of stochastic differential equations
G. N. Mil’shtein · 1975
Earlier work this paper cites.
On solving certain nonlinear partial differential equations by accretive operator methods
L. Evans · 1980
Earlier work this paper cites.
Laplace’s method revisited: weak convergence of probability measures
C.-R. Hwang · 1980
Earlier work this paper cites.
Generalized Solutions of Hamilton-Jacobi Equations
P.-L. Lions · 1982
Earlier work this paper cites.
Analyse numérique des équations différentielles stochastiques
D. Talay · 1982
Earlier work this paper cites.
Viscosity solutions of Hamilton-Jacobi equations
M. Crandall and P.-L. Lions · 1983
Earlier work this paper cites.
Some properties of viscosity solutions of Hamilton-Jacobi equations
M. Crandall, L. Evans, and P.-L. Lions · 1984
Earlier work this paper cites.
Efficient numerical schemes for the approximation of expectations of functionals of the solution of a SDE and applications
D. Talay · 1984
Earlier work this paper cites.
Discretization and simulation of stochastic differential equations
E. Pardoux and D. Talay · 1985
Earlier work this paper cites.
Weak approximation of solutions of systems of stochastic differential equations
G. N. Mil’shtein · 1986
Earlier work this paper cites.
A Mathematical Introduction to Fluid Mechanics
A. Chorin and J. Marsden · 1990
Earlier work this paper cites.
The approximation of multiple stochastic integrals
P. E. Kloeden and E. Platen · 1992
Earlier work this paper cites.
Dynamics of Langevin simulations
A. S. Kronfeld · 1993
Earlier work this paper cites.
The law of the Euler scheme for stochastic differential equations: Ii. convergence rate of the density
V. Bally and D. Talay · 1996
Earlier work this paper cites.
Topological Methods in Hydrodynamics
V. I. Arnol’d and B. A. Khesin · 1999
Earlier work this paper cites.
Vanishing viscosity limit for initial-boundary value problems for conservation laws
G.-Q. Chen and H. Frid · 1999
Earlier work this paper cites.
On the trend to equilibrium for the Fokker-Planck equation: An interplay between physics and functional analysis
P. A. Markowich and C. Villani · 1999
Earlier work this paper cites.
Large deviations
J.-D. Deuschel and D. Stroock · 2001
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications
H. Kushner and G. G. Yin · 2003
Earlier work this paper cites.
Semiconcave Functions, Hamilton-Jacobi Equations, and Optimal Control
P. Cannarsa and C. Sinestrari · 2004
Earlier work this paper cites.
Quantitative analysis of metastability in reversible diffusion processes via a Witten complex approach
B. Helffer, M. Klein, and F. Nier · 2004
Earlier work this paper cites.
Quantitative analysis of metastability in reversible diffusion processes via a Witten complex approach
F. Nier · 2004
Earlier work this paper cites.
Metastability in reversible diffusion processes II: Precise asymptotics for small eigenvalues
A. Bovier, V. Gayrard, and M. Klein · 2005
Earlier work this paper cites.
Hypoelliptic estimates and spectral theory for Fokker-Planck operators and Witten Laplacians
B. Helffer and F. Nier · 2005
Earlier work this paper cites.
Hypocoercive diffusion operators
C. Villani · 2006
Earlier work this paper cites.
Quantum Physics
S. Gasiorowicz · 2007
Earlier work this paper cites.
Adaptive online gradient descent
E. Hazan, A. Rakhlin, and P. Bartlett · 2008
Earlier work this paper cites.
Fluid Mechanics (Fourth Edition)
P. Kundu, I. Cohen, and D. Dowling · 2008
Cited alongside, same era.
Learning multiple layers of features from tiny images
A Krizhevsky · 2009
Cited alongside, same era.
Hypocoercivity
C. Villani · 2009
Cited alongside, same era.
Large-scale machine learning with stochastic gradient descent
L. Bottou · 2010
Cited alongside, same era.
Partial Differential Equations (Second Edition)
L. Evans · 2010
Cited alongside, same era.
Tunnel effect and symmetries for Kramers–Fokker–Planck type operators
F. Hérau, M. Hitrik, and J. Sjöstrand · 2011
Cited alongside, same era.
Geometrical Methods in the Theory of Ordinary Differential Equations
How to escape saddle points efficiently
C. Jin, R. Ge, P. Netrapalli, S. M. Kakade, and M. I. Jordan · 2017
Later among the works it cites.
Three factors influencing minima in SGD
S. Jastrzebski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2017
Later among the works it cites.
Acceleration and averaging in stochastic descent dynamics
W. Krichene and P. L. Bartlett · 2017
Later among the works it cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Q. Li, C. Tai, and W. E · 2017
Later among the works it cites.
Non-convex learning via stochastic gradient Langevin dynamics: A nonasymptotic analysis
M. Raginsky, A. Rakhlin, and M. Telgarsky · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Arnol’d · 2012
Cited alongside, same era.
Practical recommendations for gradient-based training of deep architectures
Y. Bengio · 2012
Cited alongside, same era.
An Introduction to Stochastic Differential Equations
L. Evans · 2012
Cited alongside, same era.
Random Perturbations of Dynamical Systems
M. Freidlin and A. Wentzell · 2012
Cited alongside, same era.
Introduction to Spectral Theory: With Applications to Schrödinger Operators
P. Hislop and I. Sigal · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
S. L. Smith, P.-J. Kindermans, C. Ying, and Q. V. Le · 2017
Later among the works it cites.
Cyclical learning rates for training neural networks
L. N. Smith · 2017
Later among the works it cites.
A hitting time analysis of stochastic gradient Langevin dynamics
Y. Zhang, P. Liang, and M. Charikar · 2017
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Later among the works it cites.
Deep relaxation: partial differential equations for optimizing deep neural networks
P. Chaudhari, A. Oberman, S. Osher, S. Soatto, and G. Carlier · 2018
Later among the works it cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
P. Chaudhari and S. Soatto · 2018
Later among the works it cites.
Characterizing implicit bias in terms of optimization geometry
S. Gunasekar, J. Lee, D. Soudry, and N. Srebro · 2018
Later among the works it cites.
On the relation between the sharpest directions of DNN loss and the SGD step length
S. Jastrzebski, Z. Kenton, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2018
Later among the works it cites.
Dynamical, symplectic and stochastic perspectives on gradient-based optimization
M. I. Jordan · 2018
Later among the works it cites.
Understanding the acceleration phenomenon via high-resolution differential equations
B. Shi, S. Du, M. Jordan, and W. J. Su · 2018
Later among the works it cites.
Gradient flow algorithms for density propagation in stochastic systems
K. Caluya and A. Halder · 2019
Later among the works it cites.
Stochastic algorithms with geometric step decay converge linearly on sharp functions
D. Davis, D. Drusvyatskiy, and V. Charisopoulos · 2019
Later among the works it cites.
Generalized momentum-based methods: A Hamiltonian perspective
J. Diakonikolas and M. I. Jordan · 2019
Later among the works it cites.
An exponential learning rate schedule for deep learning
Z. Li and S. Arora · 2019
Later among the works it cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Y. Li, C. Wei, and T. Ma · 2019
Later among the works it cites.
About small eigenvalues of Witten Laplacian
L. Michel · 2019
Later among the works it cites.
Robust learning rate selection for stochastic optimization via splitting diagnostic
M. Sordello and W. J. Su · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
U. Simsekli, L. Sagun, and M. Gurbuzbalaban · 2019
Later among the works it cites.
Optimization for deep learning: theory and algorithms
R. Sun · 2019
Later among the works it cites.
How does learning rate decay help modern neural networks?
K. You, M. Long, J. Wang, and M. I. Jordan · 2019
Later among the works it cites.
The local elasticity of neural networks
H. He and W. J. Su · 2020
Closest in time.