Fetching the paper…
Reading the bibliography…
Momentum Stochastic Gradient Descent (MSGD) algorithm has been widely applied to many nonconvex optimization problems in machine learning, e.g., training deep neural networks, variational Bayesian inference, and etc.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Mei, S · 1902
Earlier work this paper cites.
Mean field analysis of deep neural networks
Sirignano, J · 1903
Earlier work this paper cites.
A stochastic approximation method
Robbins, H · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T · 1964
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate O ( 1 / k 2 ) {O}(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
Stochastic approximation methods for systems over an infinite horizon
Kushner, H. J · 1996
Earlier work this paper cites.
Stochastic approximation with two time scales
Borkar, V. S · 1997
Earlier work this paper cites.
Brownian motion
Karatzas, I · 1998
Earlier work this paper cites.
The ODE method for convergence of stochastic approximation and reinforcement learning
Borkar, V. S · 2000
Earlier work this paper cites.
Stochastic Approximation and Recursive Algorithms and Applications
Kushner, H. J · 2003
Earlier work this paper cites.
Stochastic differential equations
Øksendal, B · 2003
Earlier work this paper cites.
Theory of ordinary differential equations: Existence, uniqueness and stability
Hu, J · 2004
Earlier work this paper cites.
Stochastic Approximation: A Dynamical Systems Viewpoint
Borkar, V. S · 2009
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S · 2013
Earlier work this paper cites.
On Multi-parameter Semimartingales, Their Integrals and Weak Convergence
Nowakowski, B. D · 2013
Earlier work this paper cites.
Weak convergence of probability measures
Sagitov, S · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P · 2014
Cited alongside, same era.
The loss surfaces of multilayer networks
Choromanska, A · 2015
Cited alongside, same era.
Handbook of Simulation Optimization: An Overview of Stochastic Approximation
Fu, M. C · 2015
Cited alongside, same era.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Abadi, M · 2016
Cited alongside, same era.
Matrix completion has no spurious local minimum
Ge, R · 2016
Acceleration and averaging in stochastic mirror descent dynamics
Krichene, W · 2017
Later among the works it cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Li, Q · 2017
Later among the works it cites.
Exploring generalization in deep learning
Neyshabur, B · 2017
Later among the works it cites.
Wang, Y · 2017
Later among the works it cites.
Theory of deep learning iii: Generalization properties of SGD
Zhang, C · 2017
Later among the works it cites.
Dimensionality reduction for stationary time series via stochastic nonconvex optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
Ghadimi, S · 2016
Cited alongside, same era.
Deep learning
Goodfellow, I · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S · 2016
Cited alongside, same era.
Symmetry, saddle points, and global geometry of nonconvex matrix factorization
Li, X · 2016
Cited alongside, same era.
Differential equations with applications and historical notes
Simmons, G. F · 2016
Cited alongside, same era.
Chen, M · 2018
Closest in time.
Liu, T · 2018
Closest in time.
Gaussian process behaviour in wide deep neural networks
Matthews, A. G. d. G · 2018
Closest in time.
A mean field view of the landscape of two-layer neural networks
Mei, S · 2018
Closest in time.
Recent trends in stochastic gradient descent for machine learning and big data
Newton, D · 2018
Closest in time.
How to train your ResNet
Page, D · 2018
Closest in time.
Neural networks as interacting particle systems: Asymptotic convexity of the loss landscape and universal scaling of the approximation error
Rotskoff, G. M · 2018
Closest in time.
Mean field analysis of neural networks
Sirignano, J · 2018
Closest in time.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A · 2019
Closest in time.
Toward understanding the importance of noise in training neural networks
Zhou, M · 2019
Closest in time.