Fetching the paper…
Reading the bibliography…
A theoretical, and potentially also practical, problem with stochastic gradient descent is that trajectories may escape to infinity.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
Analysis of recursive stochastic algorithms
L. Ljung · 1977
Earlier work this paper cites.
Asymptotic behavior of dissipative systems
J. K. Hale · 1988
Earlier work this paper cites.
A simple weight decay can improve generalization
A. Krogh and J. A. Hertz · 1992
Earlier work this paper cites.
Ergodicity for sdes and approximations: locally lipschitz vector fields and degenerate noise
J. C. Mattingly, A. M. Stuart, and D. J. Higham · 2002
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
E. Moulines and F. R. Bach · 2011
Earlier work this paper cites.
Robust penalized logistic regression with truncated loss functions
S. Y. Park and Y. Liu · 2011
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
E. Hazan and S. Kale · 2014
Earlier work this paper cites.
Guaranteed matrix completion via non-convex factorization
R. Sun and Z.-Q. Luo · 2016
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2017
Cited alongside, same era.
Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis
M. Raginsky, A. Rakhlin, and M. Telgarsky · 2017
Cited alongside, same era.
Global non-convex optimization with discretized diffusions
M. A. Erdogdu, L. Mackey, and O. Shamir · 2018
Cited alongside, same era.
The step decay schedule: A near optimal, geometrically decaying learning rate procedure for least squares
R. Ge, S. M. Kakade, R. Kidambi, and P. Netrapalli · 2019
Cited alongside, same era.
A diffusion process perspective on posterior contraction rates for parameters
W. Mou, N. Ho, M. J. Wainwright, P. Bartlett, and M. I. Jordan · 2019
Cited alongside, same era.
Stationary behavior of constant stepsize sgd type algorithms: An asymptotic characterization
Z. Chen, S. Mou, and S. T. Maguluri · 2021
Later among the works it cites.
On the convergence of langevin monte carlo: The interplay between tail growth and smoothness
M. A. Erdogdu and R. Hosseinzadeh · 2021
Later among the works it cites.
A second look at exponential and cosine step sizes: Simplicity, adaptivity, and performance
X. Li, Z. Zhuang, and F. Orabona · 2021
Later among the works it cites.
Stochastic gradient descent on nonconvex functions with general noise models
V. Patel and S. Zhang · 2021
Later among the works it cites.
Almost sure convergence rates for stochastic gradient descent and stochastic heavy ball
O. Sebbouh, R. M. Gower, and A. Defazio · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. S. Tan and R. Vershynin · 2019
Cited alongside, same era.
Stagewise training accelerates convergence of testing error over sgd
Z. Yuan, Y. Yan, R. Jin, and T. Yang · 2019
Cited alongside, same era.
Stochastic subgradient method converges on tame functions
D. Davis, D. Drusvyatskiy, S. Kakade, and J. D. Lee · 2020
Cited alongside, same era.
On the almost sure convergence of stochastic gradient descent in non-convex problems
P. Mertikopoulos, N. Hallak, A. Kavis, and V. Cevher · 2020
Cited alongside, same era.
On learning rates and schr \ \backslash " odinger operators
B. Shi, W. J. Su, and M. I. Jordan · 2020
Cited alongside, same era.
Learning with non-convex truncated losses by sgd
Y. Xu, S. Zhu, S. Yang, C. Zhang, R. Jin, and T. Yang · 2020
Cited alongside, same era.
Stochastic optimization for dc functions and non-smooth non-convex regularizers with non-asymptotic convergence
Y. Xu, Q. Qi, Q. Lin, R. Jin, and T. Yang
Cited in the paper.
Bandwidth-based step-sizes for non-convex stochastic optimization
X. Wang and M. Johansson · 2021
Later among the works it cites.
On the convergence of stochastic gradient descent with bandwidth-based step size
X. Wang and Y.-x. Yuan · 2021
Later among the works it cites.
On the convergence of step decay step-size for stochastic optimization
X. Wang, S. Magnússon, and M. Johansson · 2021
Later among the works it cites.
Stochastic gradient descent with noise of machine learning type. part i: Discrete time analysis
S. Wojtowytsch · 2021
Later among the works it cites.
An analysis of constant step size sgd in the non-convex regime: Asymptotic normality and bias
L. Yu, K. Balasubramanian, S. Volgushev, and M. A. Erdogdu · 2021
Later among the works it cites.