Fetching the paper…
Reading the bibliography…
Rapid advances in data collection and processing capabilities have allowed for the use of increasingly complex models that give rise to nonconvex optimization problems.
“Some NP-complete problems in quadratic and nonlinear programming,”
K. G. Murty and Santosh N. Kabadi, · 1987
Earlier work this paper cites.
“Recursive stochastic algorithms for global optimization in ℝ d \mathbb{R}^{d} ,”
S. Gelfand and S. Mitter, · 1991
Earlier work this paper cites.
Introductory Lectures on Convex Programming Volume I: Basic Course
Y. Nesterov, · 1998
Earlier work this paper cites.
“Gradient convergence in gradient methods with errors,”
D. Bertsekas and J. Tsitsiklis, · 2000
Earlier work this paper cites.
Matrix Analysis
R. A. Horn and C. R. Johnson, · 2003
Earlier work this paper cites.
“On the learning behavior of adaptive networks – Part II: Performance analysis,”
J. Chen and A. H. Sayed, · 2003
Earlier work this paper cites.
Convex Optimization
S. Boyd and L. Vandenberghe, · 2004
Earlier work this paper cites.
“The Perron-Frobenius theorem: Some of its applications,”
S. U. Pillai, T. Suel, and S. Cha, · 2005
Earlier work this paper cites.
“Cubic regularization of newton method and its global performance,”
Y. Nesterov and B.T. Polyak, · 2006
Earlier work this paper cites.
“Matrix factorization techniques for recommender systems,”
Y. Koren, R. Bell, and C. Volinsky, · 2009
Earlier work this paper cites.
“Efficient online and batch learning using forward backward splitting,”
J. Duchi and Y. Singer, · 2009
Earlier work this paper cites.
“Distributed subgradient methods for multi-agent optimization,”
A. Nedic and A. Ozdaglar, · 2009
Earlier work this paper cites.
“Parallelized stochastic gradient descent,”
M. Zinkevich, M. Weimer, L. Li, and A. J. Smola, · 2010
Earlier work this paper cites.
“Constrained consensus and optimization in multi-agent networks,”
A. Nedic, A. Ozdaglar, and P. A. Parrilo, · 2010
Earlier work this paper cites.
“Dual averaging for distributed optimization: Convergence analysis and network scaling,”
J. C. Duchi, A. Agarwal, and M. J. Wainwright, · 2012
Earlier work this paper cites.
“Distributed dual averaging for convex optimization under communication delays,”
K. I. Tsianos and M. G. Rabbat, · 2012
Earlier work this paper cites.
“Distributed Pareto optimization via diffusion strategies,”
J. Chen and A. H. Sayed, · 2013
Earlier work this paper cites.
“Adaptation, learning, and optimization over networks,”
A. H. Sayed, · 2014
Earlier work this paper cites.
“Adaptive networks,”
A. H. Sayed, · 2014
Earlier work this paper cites.
“On the linear convergence of the ADMM in decentralized consensus optimization,”
W. Shi, Q. Ling, K. Yuan, G. Wu, and W. Yin, · 2014
Cited alongside, same era.
“Communication-efficient distributed dual coordinate ascent,”
M. Jaggi, V. Smith, M. Takáč, J. Terhorst, S. Krishnan, T. Hofmann, and M. I. Jordan, · 2014
Cited alongside, same era.
“Deep learning,”
Y. LeCun, Y. Bengio, and G. Hinton, · 2015
Cited alongside, same era.
“The loss surfaces of multilayer networks,”
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun, · 2015
Cited alongside, same era.
“Escaping from saddle points—online stochastic gradient for tensor decomposition,”
R. Ge, F. Huang, C. Jin, and Y. Yuan, · 2015
Cited alongside, same era.
“On the learning behavior of adaptive networks - Part I: Transient analysis,”
J. Chen and A. H. Sayed, · 2015
Cited alongside, same era.
“Communication-efficient learning of deep networks from decentralized data,”
H Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, et al., · 2017
Later among the works it cites.
“ d 2 d^{2} : Decentralized training over decentralized data,”
H. Tang, X. Lian, M. Yan, C. Zhang, and J. Liu, · 2018
Later among the works it cites.
“Stochastic cubic regularization for fast nonconvex optimization,”
N. Tripuraneni, M. Stern, C. Jin, J. Regier, and M. I. Jordan, · 2018
Later among the works it cites.
“SPIDER: Near-optimal non-convex optimization via stochastic path-integrated differential estimator,”
C. Fang, C. J. Li, Z. Lin, and T. Zhang, · 2018
Later among the works it cites.
“NEON2: Finding local minima via first-order oracles,”
Z. Allen-Zhu and Y. Li, · 2018
Later among the works it cites.
“Natasha 2: Faster non-convex optimization than SGD,”
Z. Allen-Zhu, · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“DLM: Decentralized linearized alternating direction method of multipliers,”
Q. Ling, W. Shi, G. Wu, and A. Ribeiro, · 2015
Cited alongside, same era.
“Linear convergence rate of a class of distributed augmented Lagrangian algorithms,”
D. Jakovetić, J. M. F. Moura, and J. Xavier, · 2015
Cited alongside, same era.
“Stochastic variance reduction for nonconvex optimization,”
S. J. Reddi, A. Hefny, S. Sra, B. Póczós, and A. Smola, · 2016
Cited alongside, same era.
“Next: In-network nonconvex optimization,”
P. Di Lorenzo and G. Scutari, · 2016
Cited alongside, same era.
“Gradient descent only converges to minimizers,”
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht, · 2016
Cited alongside, same era.
“Matrix completion has no spurious local minimum,”
R. Ge, J. D. Lee, and T. Ma, · 2016
Cited alongside, same era.
Later among the works it cites.
“Second-order guarantees of distributed gradient algorithms,”
A. Daneshmand, G. Scutari and V. Kungurtsev, · 2018
Later among the works it cites.
“Escaping saddles with stochastic gradients,”
H. Daneshmand, J. Kohler, A. Lucchi and T. Hofmann, · 2018
Later among the works it cites.
“A linear algorithm for optimization over directed graphs with geometric convergence,”
R. Xin and U. A. Khan, · 2018
Later among the works it cites.
“Global convergence of ADMM in nonconvex nonsmooth optimization,”
Y. Wang, W. Yin, and J. Zeng, · 2019
Later among the works it cites.
“Sharp analysis for nonconvex SGD escaping from saddle points,”
C. Fang, Z. Lin and T. Zhang, · 2019
Later among the works it cites.
“Stochastic gradient descent escapes saddle points efficiently,”
C. Jin, P. Netrapalli, R. Ge, S. M. Kakade and M. I. Jordan, · 2019
Later among the works it cites.
“Annealing for distributed global optimization,”
B. Swenson, S. Kar, H. V. Poor and J. M. F. Moura, · 2019
Later among the works it cites.
“Second-order guarantees of stochastic gradient descent in non-convex optimization,”
S. Vlaski and A. H. Sayed, · 2019
Later among the works it cites.
“Distributed learning in non-convex environments – Part I: Agreement at a linear rate,”
S. Vlaski and A. H. Sayed, · 2019
Later among the works it cites.
“Distributed learning in non-convex environments – Part II: Polynomial escape from saddle-points,”
S. Vlaski and A. H. Sayed, · 2019
Later among the works it cites.
“Exact diffusion for distributed optimization and learning—Part I: Algorithm development,”
K. Yuan, B. Ying, X. Zhao, and A. H. Sayed, · 2019
Later among the works it cites.
“A unification and generalization of exact distributed first-order methods,”
D. Jakovetić, · 2019
Later among the works it cites.
“Linear speedup in saddle-point escape for decentralized non-convex optimization,”
S. Vlaski and A. H. Sayed, · 2019
Later among the works it cites.