Fetching the paper…
Reading the bibliography…
In centralized settings, it is well known that stochastic gradient descent (SGD) avoids saddle points and converges to local minima in nonconvex problems.
Theory of Ordinary Differential Equations
Earl A Coddington and Norman Levinson · 1955
Earlier work this paper cites.
Differential Inclusions: Set-Valued Maps and Viability Theory
J-P Aubin and Arrigo Cellina · 1984
Earlier work this paper cites.
Optimization and Nonsmooth Analysis
Frank H Clarke · 1990
Earlier work this paper cites.
Nonconvergence to unstable points in urn models and stochastic approximations
Robin Pemantle · 1990
Earlier work this paper cites.
Probability with martingales
David Williams · 1991
Earlier work this paper cites.
A dynamical system approach to stochastic approximations
Michel Benaim · 1996
Earlier work this paper cites.
Spectral Graph Theory
Fan R. K. Chung · 1997
Earlier work this paper cites.
Distributed optimization in sensor networks
Michael Rabbat and Robert Nowak · 2004
Earlier work this paper cites.
Stochastic approximations and differential inclusions
Michel Benaïm, Josef Hofbauer, and Sylvain Sorin · 2005
Earlier work this paper cites.
Ordinary Differential Equations with Applications , volume 34 of Texts in Applied Mathematics
Carmen Chicone · 2006
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
Angelia Nedic and Asuman Ozdaglar · 2009
Earlier work this paper cites.
Gossip algorithms for distributed signal processing
Alexandros G Dimakis, Soummya Kar, José MF Moura, Michael G Rabbat, and Anna Scaglione · 2010
Earlier work this paper cites.
Probability: Theory and Examples , volume 49 of Cambridge Series in Statistical and Probabilistic Mathematics
Rick Durrett · 2010
Earlier work this paper cites.
Distributed stochastic subgradient projection algorithms for convex optimization
S Sundhar Ram, Angelia Nedić, and Venugopal V Veeravalli · 2010
Earlier work this paper cites.
Dual averaging for distributed optimization: Convergence analysis and network scaling
John C Duchi, Alekh Agarwal, and Martin J Wainwright · 2011
Earlier work this paper cites.
Convergence of a multi-agent projected stochastic gradient algorithm for non-convex optimization
Pascal Bianchi and Jérémie Jakubowicz · 2012
Earlier work this paper cites.
Diffusion adaptation strategies for distributed optimization and learning over networks
Jianshu Chen and Ali H Sayed · 2012
Earlier work this paper cites.
Distributed parameter estimation in sensor networks: Nonlinear observation models and imperfect communication
Soummya Kar, José MF Moura, and Kavita Ramanan · 2012
Earlier work this paper cites.
Distributed linear parameter estimation: Asymptotically efficient adaptive strategies
Soummya Kar, José MF Moura, and H Vincent Poor · 2013
Earlier work this paper cites.
Perturbation Theory for Linear Operators
Tosio Kato · 2013
Earlier work this paper cites.
Analysis 2
Konrad Königsberger · 2013
Earlier work this paper cites.
D-ADMM: A communication-efficient distributed algorithm for separable optimization
Joao FC Mota, Joao MF Xavier, Pedro MQ Aguiar, and Markus Püschel · 2013
Earlier work this paper cites.
Global Stability of Dynamical Systems
Michael Shub · 2013
Earlier work this paper cites.
Fast distributed gradient methods
Dušan Jakovetić, Joao Xavier, and José MF Moura · 2014
Cited alongside, same era.
Distributed optimization over time-varying directed graphs
Angelia Nedić and Alex Olshevsky · 2014
Cited alongside, same era.
A differential equation for modeling Nesterov’s accelerated gradient method: Theory and insights
Weijie Su, Stephen Boyd, and Emmanuel Candes · 2014
Cited alongside, same era.
Measure theory and fine properties of functions
Lawrence C Evans and Ronald F Gariepy · 2015
Cited alongside, same era.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Cited alongside, same era.
A unification and generalization of exact distributed first-order methods
Dušan Jakovetić · 2018
Later among the works it cites.
A distributed quasi-Newton algorithm for empirical risk minimization with nonsmooth regularization
Ching-Pei Lee, Cong Han Lim, and Stephen J Wright · 2018
Later among the works it cites.
Asy-sonata: Achieving linear convergence in distributed asynchronous multiagent optimization
Ye Tian, Ying Sun, and Gesualdo Scutari · 2018
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Later among the works it cites.
Perturbed proximal primal–dual algorithm for nonconvex nonsmooth optimization
Davood Hajinezhad and Mingyi Hong · 2019
Later among the works it cites.
Decentralized stochastic optimization and gossip algorithms with compressed communication
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Federated optimization: Distributed optimization beyond the datacenter
Jakub Konečnỳ, Brendan McMahan, and Daniel Ramage · 2015
Cited alongside, same era.
Accelerated mirror descent in continuous and discrete time
Walid Krichene, Alexandre Bayen, and Peter L Bartlett · 2015
Cited alongside, same era.
Efficient approaches for escaping higher order saddle points in non-convex optimization
Animashree Anandkumar and Rong Ge · 2016
Cited alongside, same era.
Next: In-network nonconvex optimization
Paolo Di Lorenzo and Gesualdo Scutari · 2016
Cited alongside, same era.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Cited alongside, same era.
Anastasia Koloskova, Sebastian U Stich, and Martin Jaggi · 2019
Later among the works it cites.
Distributed stochastic nonsmooth nonconvex optimization
Vyacheslav Kungurtsev · 2019
Later among the works it cites.
First-order methods almost always avoid strict saddle points
Jason D Lee, Ioannis Panageas, Georgios Piliouras, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2019
Later among the works it cites.
Revisiting normalized gradient descent: Fast evasion of saddle points
Ryan Murray, Brian Swenson, and Soummya Kar · 2019
Later among the works it cites.
Optimal convergence rates for convex distributed optimization in networks
Kevin Scaman, Francis Bach, Sébastien Bubeck, Yin Lee, and Laurent Massoulié · 2019
Later among the works it cites.
Distributed nonconvex constrained optimization over time-varying digraphs
Gesualdo Scutari and Ying Sun · 2019
Later among the works it cites.
Distributed non-convex first-order optimization and information processing: Lower complexity bounds and rate optimal algorithms
Haoran Sun and Mingyi Hong · 2019
Later among the works it cites.
Adaptive federated learning in resource constrained edge computing systems
Shiqiang Wang, Tiffany Tuor, Theodoros Salonidis, Kin K Leung, Christian Makaya, Ting He, and Kevin Chan · 2019
Later among the works it cites.
Distributed learning in the nonconvex world: From batch data to streaming and beyond
Tsung-Hui Chang, Mingyi Hong, Hoi-To Wai, Xinwei Zhang, and Songtao Lu · 2020
Closest in time.
Second-order guarantees of distributed gradient algorithms
Amir Daneshmand, Gesualdo Scutari, and Vyacheslav Kungurtsev · 2020
Closest in time.
Pathological subgradient dynamics
Aris Daniilidis and Dmitriy Drusvyatskiy · 2020
Closest in time.
Stochastic subgradient method converges on tame functions
Damek Davis, Dmitriy Drusvyatskiy, Sham Kakade, and Jason D Lee · 2020
Closest in time.
Understanding notions of stationarity in nonsmooth optimization: A guided tour of various constructions of subdifferential for nonsmooth functions
Jiajin Li, Anthony Man-Cho So, and Wing-Kin Ma · 2020
Closest in time.
On distributed stochastic gradient algorithms for global optimization
Brian Swenson, Anirudh Sridhar, and H Vincent Poor · 2020
Closest in time.
A dual approach for optimal algorithms in distributed optimization over networks
César A Uribe, Soomin Lee, Alexander Gasnikov, and Angelia Nedić · 2020
Closest in time.
On nonconvex optimization for machine learning: Gradients, stochasticity, and saddle points
Chi Jin, Praneeth Netrapalli, Rong Ge, Sham M Kakade, and Michael I Jordan · 2021
Closest in time.
Understanding the acceleration phenomenon via high-resolution differential equations
Bin Shi, Simon S Du, Michael I Jordan, and Weijie J Su · 2021
Closest in time.
Distributed gradient flow: Nonsmoothness, nonconvexity, and saddle point evasion
Brian Swenson, Ryan Murray, H Vincent Poor, and Soummya Kar · 2021
Closest in time.