Fetching the paper…
Reading the bibliography…
Stochastic gradient descent (SGD) has been found to be surprisingly effective in training a variety of deep neural networks.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Taylor expansion of the accumulated rounding error
S. Linnainmaa · 1976
Earlier work this paper cites.
Training a 3-node neural network is NP-complete
A. Blum and R. L. Rivest · 1988
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
E. Moulines and F. R Bach · 2011
Earlier work this paper cites.
An optimal method for stochastic composite optimization
G. Lan · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm
D. Needell, R. Ward, and N. Srebro · 2014
Cited alongside, same era.
Escaping from saddle points — online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Cited alongside, same era.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2016
Cited alongside, same era.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
S. Ghadimi and G. Lan · 2016
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2017
Later among the works it cites.
Convergence analysis of proximal gradient with momentum for nonconvex optimization
Q. Li, Y. Zhou, Y. Liang, and P. K. Varshney · 2017
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
M. Schmidt, N. Le Roux, and F. Bach · 2017
Later among the works it cites.
Recovery guarantees for one-hidden-layer neural networks
K. Zhong, Z. Song, P. Jain, P. L. Bartlett, and I. S. Dhillon · 2017
Later among the works it cites.
Characterization of gradient dominance and regularity conditions for neural networks
Y. Zhou and Y. Liang · 2017
Later among the works it cites.
Escaping saddles with stochastic gradients
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
S. Ghadimi, G. Lan, and H. Zhang · 2016
Cited alongside, same era.
Low-rank solutions of linear matrix equations via Procrustes flow
S. Tu, R. Boczar, M. Simchowitz, M. Soltanolkotabi, and B. Recht · 2016
Cited alongside, same era.
Geometrical properties and accelerated gradient solvers of non-convex phase retrieval
Y. Zhou, H. Zhang, and Y. Liang · 2016
Cited alongside, same era.
How to escape saddle points efficiently
C. Jin, R. Ge, P. Netrapalli, S. M. Kakade, and M. I. Jordan · 2017
Cited alongside, same era.
A generic approach for escaping saddle points
S. Reddi, M. Zaheer, S. Sra, B. Poczos, F. Bach, R. Salakhutdinov, and A. Smola
Cited in the paper.
On the convergence of Adam and beyond
S. J. Reddi, S. Kale, and S. Kumar
Cited in the paper.
H. Daneshmand, J. M. Kohler, A. Lucchi, and T. Hofmann · 2018
Later among the works it cites.
An alternative view: When does SGD escape local minima?
B. Kleinberg, Y. Li, and Y. Yuan · 2018
Later among the works it cites.
Rapid, robust, and reliable blind deconvolution via nonconvex optimization
X. Li, S. Ling, T. Strohmer, and K. Wei · 2018
Later among the works it cites.
SpiderBoost: A class of faster variance-reduced algorithms for nonconvex optimization
Z. Wang, K. Ji, Y. Zhou, Y. Liang, and V. Tarokh · 2018
Later among the works it cites.
Critical points of linear neural networks: Analytical forms and landscape properties
Y. Zhou and Y. Liang · 2018
Later among the works it cites.