Fetching the paper…
Reading the bibliography…
The article discusses distributed gradient-descent algorithms for computing local and global minima in nonconvex optimization.
1903
Earlier work this paper cites.
X. Chen, S. Du, and X. Tong, “On stationary-point hitting time and ergodicity of stochastic gradient Langevin dynamics,” 2019, online: https://arxiv.org/pdf/1904:13016.pdf
1904
Earlier work this paper cites.
1907
Earlier work this paper cites.
1907
Earlier work this paper cites.
1907
Earlier work this paper cites.
1910
Earlier work this paper cites.
E. A. Coddington and N. Levinson, Theory of Ordinary Differential Equations . Tata McGraw-Hill, 1955
1955
Earlier work this paper cites.
C.-R. Hwang et al. , “Laplace’s method revisited: Weak convergence of probability measures,” The Annals of Probability , vol. 8, no. 6, pp. 1177–1182, 1980
1980
Earlier work this paper cites.
S. Kirkpatrick, C. D. Gelatt, and M. P. Vecchi, “Optimization by simulated annealing,” Science , vol. 220, no. 4598, pp. 671–680, 1983
1983
Earlier work this paper cites.
H. J. Kushner, “Asymptotic global behavior for stochastic approximation and diffusions with slowly decreasing noise effects: Global minimization via Monte Carlo,” SIAM Journal on Applied Mathematics , vol. 47, no. 1, pp. 169–185, 1987
1987
Earlier work this paper cites.
T.-S. Chiang, C.-R. Hwang, and S. J. Sheu, “Diffusion for global optimization in ℝ n \mathbb{R}^{n} ,” SIAM Journal on Control and Optimization , vol. 25, no. 3, pp. 737–753, 1987
1987
Earlier work this paper cites.
R. Pemantle, “Nonconvergence to unstable points in urn models and stochastic approximations,” The Annals of Probability , vol. 18, no. 2, pp. 698–712, 1990
1990
Earlier work this paper cites.
S. B. Gelfand and S. K. Mitter, “Recursive stochastic algorithms for global optimization in ℝ d \mathbb{R}^{d} ,” SIAM Journal on Control and Optimization , vol. 29, no. 5, pp. 999–1018, 1991
1991
Earlier work this paper cites.
G. Yin, “Rates of convergence for a class of global stochastic optimization algorithms,” SIAM Journal on Optimization , vol. 10, no. 1, pp. 99–120, 1999
1999
Earlier work this paper cites.
M. Benaïm, “Dynamics of stochastic approximation algorithms,” in Seminaire de Probabilites 33 . Springer, 1999, pp. 1–68
1999
Earlier work this paper cites.
2003
Cited alongside, same era.
M. Rabbat and R. Nowak, “Distributed optimization in sensor networks,” in Proceedings of the 3rd international symposium on information processing in sensor networks , 2004, pp. 20–27
2004
Cited alongside, same era.
R. Olfati-Saber, J. A. Fax, and R. M. Murray, “Consensus and cooperation in networked multi-agent systems,” Proceedings of the IEEE , vol. 95, no. 1, pp. 215–233, 2007
2007
Cited alongside, same era.
2008
Cited alongside, same era.
K. Kawaguchi, “Deep learning without poor local minima,” in Proceedings of Advances in Neural Information Processing Systems , 2016, pp. 586–594
2016
Later among the works it cites.
R. Ge, J. D. Lee, and T. Ma, “Matrix completion has no spurious local minimum,” in Advances in Neural Information Processing Systems , 2016, pp. 2973–2981
2016
Later among the works it cites.
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht, “Gradient descent only converges to minimizers,” in Conference on Learning Theory, PMLR , 2016, pp. 1246–1257
2016
Later among the works it cites.
A. Wibisono, A. C. Wilson, and M. I. Jordan, “A variational perspective on accelerated methods in optimization,” Proceedings of the National Academy of Sciences , vol. 113, no. 47, pp. E7351–E7358, 2016
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Bullo, J. Cortes, and S. Martinez, Distributed Control of Robotic Networks: A Mathematical Approach to Motion Coordination Algorithms . Princeton University Press, 2009
2009
Cited alongside, same era.
A. Nedić and A. Ozdaglar, “Distributed subgradient methods for multi-agent optimization,” IEEE Transactions on Automatic Control , vol. 54, no. 1, pp. 48–61, 2009
2009
Cited alongside, same era.
A. G. Dimakis, S. Kar, J. M. Moura, M. G. Rabbat, and A. Scaglione, “Gossip algorithms for distributed signal processing,” Proceedings of the IEEE , vol. 98, no. 11, pp. 1847–1864, 2010
2010
Cited alongside, same era.
M. Shub, Global Stability of Dynamical Systems . Springer Science & Business Media, 2013
2013
Cited alongside, same era.
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio, “Identifying and attacking the saddle point problem in high-dimensional non-convex optimization,” in Proceedings of Advances in Neural Information Processing Systems , 2014, pp. 2933–2941
2014
Cited alongside, same era.
W. Su, S. Boyd, and E. Candes, “A differential equation for modeling nesterov’s accelerated gradient method: Theory and insights,” in Advances in neural information processing systems , 2014, pp. 2510–2518
2014
Cited alongside, same era.
S. Shalev-Shwartz and S. Ben-David, Understanding Machine Learning: From Theory to Algorithms . Cambridge University Press, 2014
2014
Cited alongside, same era.
R. Ge, F. Huang, C. Jin, and Y. Yuan, “Escaping from saddle points—online stochastic gradient for tensor decomposition,” in Conference on Learning Theory , 2015, pp. 797–842
2015
Cited alongside, same era.
X. Lian, C. Zhang, H. Zhang, C.-J. Hsieh, W. Zhang, and J. Liu, “Can decentralized algorithms outperform centralized algorithms? A case study for decentralized parallel stochastic gradient descent,” in Proceedings of Advances in Neural Information Processing Systems , 2017, pp. 5330–5340
2017
Later among the works it cites.
C. Jin, R. Ge, P. Netrapalli, S. M. Kakade, and M. I. Jordan, “How to escape saddle points efficiently,” in Proceedings of the International Conference on Machine Learning , vol. 70, 2017, pp. 1724–1732
2017
Later among the works it cites.
M. Raginsky, A. Rakhlin, and M. Telgarsky, “Non-convex learning via stochastic gradient Langevin dynamics: A nonasymptotic analysis,” in Proceedings of Conference on Learning Theory, PMLR , 2017, pp. 1674–1703
2017
Later among the works it cites.
2017
Later among the works it cites.
2018
Later among the works it cites.
M. Hong, J. D. Lee, and M. Razaviyayn, “Gradient primal-dual algorithm converges to second-order stationary solution for nonconvex distributed optimization over networks,” in Proceedings of the International Conference on Machine Learning , 2018
2018
Later among the works it cites.
R. Murray, B. Swenson, and S. Kar, “Revisiting normalized gradient descent: Fast evasion of saddle points,” IEEE Transactions on Automatic Control , vol. 64, no. 11, pp. 4818–4824, 2019
2019
Later among the works it cites.
——, “Distributed gradient descent: Nonconvergence to saddle points and the stable-manifold theorem,” in Proceedings of Allerton Conference on Communication, Control, and Computing , 2019, pp. 595–601
2019
Later among the works it cites.
J. T. Barron, “A general and adaptive robust loss function,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 4331–4339
2019
Later among the works it cites.
D. Davis, D. Drusvyatskiy, S. Kakade, and J. D. Lee, “Stochastic subgradient method converges on tame functions,” Foundations of computational mathematics , vol. 20, no. 1, pp. 119–154, 2020
2020
Closest in time.
Y. Zhang, P. Liang, and M. Charikar, “A hitting time analysis of stochastic gradient Langevin dynamics,” in Proceedings of Conference on Learning Theory, PMLR , 2017, pp. 1980–2022
2022
Closest in time.