Fetching the paper…
Reading the bibliography…
This paper analyzes the trajectories of stochastic gradient descent (SGD) to help understand the algorithm's convergence properties in non-convex problems.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
On a stochastic approximation method
Kuo-Liang Chung · 1954
Earlier work this paper cites.
Topology from the Differentiable Viewpoint
John Willard Milnor · 1965
Earlier work this paper cites.
A convergence theorem for non negative almost supermartingales and some applications
Herbert Robbins and David Siegmund · 1971
Earlier work this paper cites.
Differential Topology
Morris W. Hirsch · 1976
Earlier work this paper cites.
Analysis of recursive stochastic algorithms
Lennart Ljung · 1977
Earlier work this paper cites.
Martingale Limit Theory and Its Application
P. Hall and C. C. Heyde · 1980
Earlier work this paper cites.
Basic Topology
Mark Anthony Armstrong · 1983
Earlier work this paper cites.
System Identification Theory for the User
Lennart Ljung · 1986
Earlier work this paper cites.
Introduction to Optimization
Boris Teodorovich Polyak · 1987
Earlier work this paper cites.
Global Stability of Dynamical Systems
Michael Shub · 1987
Earlier work this paper cites.
Adaptive Algorithms and Stochastic Approximations
Albert Benveniste, Michel Métivier, and Pierre Priouret · 1990
Earlier work this paper cites.
Nonconvergence to unstable points in urn models and stochastic aproximations
Robin Pemantle · 1990
Earlier work this paper cites.
Vertex-reinforced random walk
Robin Pemantle · 1992
Earlier work this paper cites.
Dynamics of Morse-Smale urn processes
Michel Benaïm and Morris W. Hirsch · 1995
Earlier work this paper cites.
Asymptotic pseudotrajectories and chain recurrent flows, with applications
Michel Benaïm and Morris W. Hirsch · 1996
Cited alongside, same era.
Les algorithmes stochastiques contournent-ils les pièges ?
Odile Brandière and Marie Duflo · 1996
Cited alongside, same era.
Stochastic approximation algorithms and applications
Harold J. Kushner and G. G. Yin · 1997
Cited alongside, same era.
Dynamics of stochastic approximation algorithms
Michel Benaïm · 1999
Cited alongside, same era.
Gradient convergence in gradient methods with errors
Dimitri P. Bertsekas and John N. Tsitsiklis · 2000
Cited alongside, same era.
Introduction to Smooth Manifolds
John M. Lee · 2003
Cited alongside, same era.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
Saeed Ghadimi and Guanghui Lan · 2016
Later among the works it cites.
Gradient descent only converges to minimizers
Jason D. Lee, Max Simchowitz, Michael I. Jordan, and Benjamin Recht · 2016
Later among the works it cites.
Gradient descent can take exponential time to escape saddle points
Simon S. Du, Chi Jin, Jason D. Lee, Michael I. Jordan, Barnabás Póczos, and Aarti Singh · 2017
Later among the works it cites.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Later among the works it cites.
Gradient descent only converges to minimizers: Non-isolated critical points and invariant regions
Ioannis Panageas and Georgios Piliouras · 2017
Later among the works it cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Suder, and Tom Goldstein · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Introductory Lectures on Convex Optimization: A Basic Course
Yurii Nesterov · 2004
Cited alongside, same era.
Stochastic Approximation: A Dynamical Systems Viewpoint
Vivek S. Borkar · 2008
Cited alongside, same era.
Solving variational inequalities with stochastic mirror-prox algorithm
Anatoli Juditsky, Arkadi Semen Nemirovski, and Claire Tauvel · 2011
Cited alongside, same era.
An Introduction to Dynamical Systems: Continuous and Discrete
R. Clark (Rex) Robinson · 2012
Cited alongside, same era.
Ordinary Differential Equations and Dynamical Systems , volume 140 of Graduate Studies in Mathematics
Gerald Teschl · 2012
Cited alongside, same era.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Cited alongside, same era.
Later among the works it cites.
Efficiently avoiding saddle points with zero order methods: No gradients required
Lampros Flokas, Emmanouil Vasileios Vlatakis-Gkaragkounis, and Georgios Piliouras · 2019
Later among the works it cites.
On the convergence of single-call stochastic extra-gradient methods
Yu-Guan Hsieh, Franck Iutzeler, Jérôme Malick, and Panayotis Mertikopoulos · 2019
Later among the works it cites.
First-order methods almost always avoid strict saddle points
Jason D. Lee, Ioannis Panageas, Georgios Piliouras, Max Simchowitz, Michael I. Jordan, and Benjamin Recht · 2019
Later among the works it cites.
Stochastic gradient descent for nonconvex learning without bounded gradient assumptions
Yunwen Lei, Ting Hu, Guiying Li, and Ke Tang · 2019
Later among the works it cites.
Learning in games with continuous action sets and unknown payoff functions
Panayotis Mertikopoulos and Zhengyuan Zhou · 2019
Later among the works it cites.
First-order methods almost always avoid saddle points: The case of vanishing step-sizes
Ioannis Panageas, Georgios Piliouras, and Xiao Wang · 2019
Later among the works it cites.
Second-order guarantees of stochastic gradient descent in non-convex optimization
Stefan Vlaski and Ali H. Sayed · 2019
Later among the works it cites.
Yu-Guan Hsieh, Franck Iutzeler, Jérôme Malick, and Panayotis Mertikopoulos · 2020
Closest in time.