Fetching the paper…
Reading the bibliography…
Recent years have seen increased interest in performance guarantees of gradient descent algorithms for non-convex optimization.
“A stochastic approximation method,”
H. Robbins and S. Monro, · 1951
Earlier work this paper cites.
“Recursive stochastic algorithms for global optimization in ℝ d \mathbb{R}^{d} ,”
S. Gelfand and S. Mitter, · 1991
Earlier work this paper cites.
Introduction to Optimization
B. T. Polyak, · 1997
Earlier work this paper cites.
Introductory Lectures on Convex Programming Volume I: Basic Course
Y. Nesterov, · 1998
Earlier work this paper cites.
“Gradient convergence in gradient methods with errors,”
D. Bertsekas and J. Tsitsiklis, · 2000
Earlier work this paper cites.
“Cubic regularization of newton method and its global performance,”
Y. Nesterov and B.T. Polyak, · 2006
Earlier work this paper cites.
“Distributed Pareto optimization via diffusion strategies,”
J. Chen and A. H. Sayed, · 2013
Earlier work this paper cites.
“Adaptive networks,”
A. H. Sayed, · 2014
Earlier work this paper cites.
“Adaptation, learning, and optimization over networks,”
A. H. Sayed, · 2014
Earlier work this paper cites.
“Escaping from saddle points—online stochastic gradient for tensor decomposition,”
R. Ge, F. Huang, C. Jin, and Y. Yuan, · 2015
Earlier work this paper cites.
“The Loss Surfaces of Multilayer Networks,”
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun, · 2015
Earlier work this paper cites.
“Stochastic variance reduction for nonconvex optimization,”
S. J. Reddi, A. Hefny, S. Sra, B. Póczós, and A. Smola, · 2016
Cited alongside, same era.
“Deep learning without poor local minima,”
K. Kawaguchi, · 2016
Cited alongside, same era.
“Matrix completion has no spurious local minimum,”
R. Ge, J. D. Lee, and T. Ma, · 2016
Cited alongside, same era.
“Global optimality of local search for low rank matrix recovery,”
S. Bhojanapalli, B. Neyshabur, and N. Srebro, · 2016
Cited alongside, same era.
“Gradient descent only converges to minimizers,”
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht, · 2016
Cited alongside, same era.
“Gradient descent can take exponential time to escape saddle points,”
S. S. Du, C. Jin, J. D. Lee, M. I. Jordan, B. Poczos and A. Singh, · 2017
“NEON2: Finding local minima via first-order oracles,”
Z. Allen-Zhu and Y. Li, · 2018
Later among the works it cites.
“Natasha 2: Faster non-convex optimization than SGD,”
Z. Allen-Zhu, · 2018
Later among the works it cites.
“Second-order guarantees of distributed gradient algorithms,”
A. Daneshmand, G. Scutari and V. Kungurtsev, · 2018
Later among the works it cites.
“Accelerated gradient descent escapes saddle points faster than gradient descent,”
C. Jin, P. Netrapalli, and M. I. Jordan, · 2018
Later among the works it cites.
“Escaping saddles with stochastic gradients,”
H. Daneshmand, J. Kohler, A. Lucchi and T. Hofmann, · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
F. Facchinei, V. Kungurtsev, L. Lampariello, G. Scutari, · 2017
Cited alongside, same era.
“Non-convex optimization for machine learning,”
P. Jain and P. Kar, · 2017
Cited alongside, same era.
“A trust region algorithm with a worst-case iteration complexity of o ( ϵ − 3 / 2 ) o(\epsilon^{-3/2}) for nonconvex optimization,”
F. E. Curtis, D. P. Robinson, and M. Samadi, · 2017
Cited alongside, same era.
“How to escape saddle points efficiently,”
C. Jin, R. Ge, P. Netrapalli, S. M. Kakade, and M. I. Jordan, · 2017
Cited alongside, same era.
“SPIDER: Near-optimal non-convex optimization via stochastic path-integrated differential estimator,”
C. Fang, C. J. Li, Z. Lin, and T. Zhang, · 2018
Cited alongside, same era.
C. Jin, P. Netrapalli, R. Ge, S. M. Kakade and M. I. Jordan, · 2019
Closest in time.
“Distributed learning in non-convex environments – Part I: Agreement at a linear rate,”
S. Vlaski and A. H. Sayed, · 2019
Closest in time.
“Distributed learning in non-convex environments – Part II: Polynomial escape from saddle-points,”
S. Vlaski and A. H. Sayed, · 2019
Closest in time.
“Annealing for distributed global optimization,”
B. Swenson, S. Kar, H. V. Poor and J. M. F. Moura, · 2019
Closest in time.
“Sharp analysis for nonconvex sgd escaping from saddle points,”
C. Fang, Z. Lin and T. Zhang, · 2019
Closest in time.
“Stabilized SVRG: Simple variance reduction for nonconvex optimization,”
R. Ge, Z. Li, W. Wang and X. Wang, · 2019
Closest in time.