Fetching the paper…
Reading the bibliography…
This paper considers the problem of understanding the behavior of a general class of accelerated gradient methods on smooth nonconvex functions.
Fundamenta Mathematicae 22
Kirszbraun, M.: Über die zusammenziehende und lipschitzsche transformationen · 1934
Earlier work this paper cites.
USSR Computational Mathematics and Mathematical Physics 4
Polyak, B.T.: Some methods of speeding up the convergence of iteration methods · 1964
Earlier work this paper cites.
Proceedings of the American Mathematical Society 16
Whittlesey, E.F.: Analytic functions in banach spaces · 1965
Earlier work this paper cites.
Springer (1967)
Hahn, W., et al.: Stability of motion, vol. 138 · 1967
Earlier work this paper cites.
Bulletin of the American mathematical Society 73
Smale, S.: Differentiable dynamical systems · 1967
Earlier work this paper cites.
CRC Press (1969)
Schwartz, J.T.: Nonlinear functional analysis, vol. 4 · 1969
Earlier work this paper cites.
McGraw-hill New York (1976)
Rudin, W., et al.: Principles of mathematical analysis, vol. 3 · 1976
Earlier work this paper cites.
Hirsch, M., Pugh, C., Shub, M.: Invariant manifolds (lecture notes in mathematics, 583) (1977)
1977
Earlier work this paper cites.
In: Dokl. akad. nauk Sssr, vol. 269, pp. 543–547 (1983)
Nesterov, Y.E.: A method for solving the convex programming problem with convergence rate o (1/kˆ 2) · 1983
Earlier work this paper cites.
Springer (1984)
Luenberger, D.G., Ye, Y., et al.: Linear and nonlinear programming, vol. 2 · 1984
Earlier work this paper cites.
John Wiley & Sons (1988)
Dunford, N., Schwartz, J.T.: Linear operators, part 1: general theory, vol. 10 · 1988
Earlier work this paper cites.
Wiley-Interscience (1989)
Tabor, M.: Chaos and integrability in nonlinear dynamics: an introduction · 1989
Earlier work this paper cites.
International journal of control 55
Lyapunov, A.M.: The general problem of the stability of motion · 1992
Earlier work this paper cites.
OUP Oxford (1994)
Davidson, J.: Stochastic limit theory: An introduction for econometricians · 1994
Earlier work this paper cites.
Advances in Computational mathematics 5
Corless, R.M., Gonnet, G.H., Hare, D.E., Jeffrey, D.J., Knuth, D.E.: On the lambertw function · 1996
Earlier work this paper cites.
Springer (1999)
Chicone, C.C.: Ordinary Differential Equations With Applications · 1999
Earlier work this paper cites.
Mathematics of Computation 68
Xinghua, W.: Convergence of newton’s method and inverse function theorem in banach space · 1999
Earlier work this paper cites.
SIAM (2000)
Kinderlehrer, D., Stampacchia, G.: An introduction to variational inequalities and their applications · 2000
Earlier work this paper cites.
American Mathematical Soc. (2002)
Matsumoto, Y.: An introduction to Morse theory, vol. 208 · 2002
Earlier work this paper cites.
Cambridge university press (2002)
Ott, E.: Chaos in dynamical systems · 2002
Earlier work this paper cites.
Springer Science & Business Media (2003)
Nesterov, Y.: Introductory lectures on convex optimization: A basic course, vol. 87 · 2003
Earlier work this paper cites.
Journal of Inequalities and Applications 2005
Papi, M.: On the domain of the implicit function and applications · 2005
Earlier work this paper cites.
Lecture Notes (2010)
Tibshirani, R., et al.: Proximal gradient descent and acceleration · 2010
Earlier work this paper cites.
Springer (2011)
Brezis, H., Brézis, H.: Functional analysis, Sobolev spaces and partial differential equations, vol. 2 · 2011
Earlier work this paper cites.
Springer Science & Business Media (2012)
Kirillov, A.A., Gvishiani, A.D.: Theorems and problems in functional analysis · 2012
Earlier work this paper cites.
Springer Science & Business Media (2012)
Megginson, R.E.: An introduction to Banach space theory, vol. 183 · 2012
Earlier work this paper cites.
In: Introduction to Smooth Manifolds, pp. 1–31. Springer (2013)
Lee, J.M.: Smooth manifolds · 2013
Earlier work this paper cites.
Springer Science & Business Media (2013)
Shub, M.: Global stability of dynamical systems · 2013
Earlier work this paper cites.
In: International conference on machine learning, pp. 1139–1147. PMLR (2013)
Sutskever, I., Martens, J., Dahl, G., Hinton, G.: On the importance of initialization and momentum in deep learning · 2013
Earlier work this paper cites.
Advances in neural information processing systems 27
Su, W., Boyd, S., Candes, E.: A differential equation for modeling nesterov’s accelerated gradient method: theory and insights · 2014
Earlier work this paper cites.
IEEE Transactions on Information Theory 61
Candes, E.J., Li, X., Soltanolkotabi, M.: Phase retrieval via wirtinger flow: Theory and algorithms · 2015
Earlier work this paper cites.
Bulletin of the Belgian Mathematical Society-Simon Stevin 22
Conejero, J.A., Muñoz-Fernández, G.A., Arcila, M.M., Seoane-Sepúlveda, J.B.: Smooth functions with uncountably many zeros · 2015
Cited alongside, same era.
In: 2015 European control conference (ECC), pp. 310–315. IEEE (2015)
Ghadimi, E., Feyzmahdavian, H.R., Johansson, M.: Global convergence of the heavy-ball method for convex optimization · 2015
Cited alongside, same era.
Dozat, T.: Incorporating nesterov momentum into adam (2016)
2016
Cited alongside, same era.
Mathematical Programming 156
Ghadimi, S., Lan, G.: Accelerated gradient methods for nonconvex nonlinear and stochastic programming · 2016
Cited alongside, same era.
IEEE Journal of selected topics in signal processing 10
Jaganathan, K., Eldar, Y.C., Hassibi, B.: Stft phase retrieval: Uniqueness guarantees and recovery algorithms · 2016
Cited alongside, same era.
arXiv preprint arXiv:1911.07596 (2019)
Barakat, A., Bianchi, P.: Convergence analysis of a momentum algorithm with adaptive step size for non convex optimization · 2019
Later among the works it cites.
In: International Conference on Machine Learning, pp. 891–901. PMLR (2019)
Can, B., Gurbuzbalaban, M., Zhu, L.: Accelerated linear convergence of stochastic momentum methods in wasserstein distances · 2019
Later among the works it cites.
Mathematical Programming 176
Chen, Y., Chi, Y., Fan, J., Ma, C.: Gradient descent with random initialization: Fast global convergence for nonconvex phase retrieval · 2019
Later among the works it cites.
arXiv preprint arXiv:1902.00247 (2019)
Fang, C., Lin, Z., Zhang, T.: Sharp analysis for nonconvex sgd escaping from saddle points · 2019
Later among the works it cites.
Advances in Neural Information Processing Systems 32
Gitman, I., Lang, H., Zhang, P., Xiao, L.: Understanding the role of momentum in stochastic gradient methods · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lee, J.D., Simchowitz, M., Jordan, M.I., Recht, B.: Gradient descent converges to minimizers · 2016
Cited alongside, same era.
arXiv preprint arXiv:1605.00405 (2016)
Panageas, I., Piliouras, G.: Gradient descent only converges to minimizers: Non-isolated critical points and invariant regions · 2016
Cited alongside, same era.
In: 2016 54th Annual Allerton Conference on Communication, Control, and Computing (Allerton), pp. 331–335. IEEE (2016)
Zhou, Y., Zhang, H., Liang, Y.: Geometrical properties and accelerated gradient solvers of non-convex phase retrieval · 2016
Cited alongside, same era.
Aujol, J., Dossal, C.: Optimal rate of convergence of an ode associated to the fast gradient descent schemes for b > > 0 (2017)
2017
Cited alongside, same era.
arXiv preprint arXiv:1711.10456 (2017)
Jin, C., Netrapalli, P., Jordan, M.I.: Accelerated gradient descent escapes saddle points faster than gradient descent · 2017
Cited alongside, same era.
Courier Dover Publications (2017)
Kelley, J.L.: General topology · 2017
Cited alongside, same era.
IEEE Transactions on Signal Processing 66
Pauwels, E.J.R., Beck, A., Eldar, Y.C., Sabach, S.: On fienup methods for sparse phase retrieval · 2017
Cited alongside, same era.
Later among the works it cites.
Mathematical programming 176
Lee, J.D., Panageas, I., Piliouras, G., Simchowitz, M., Jordan, M.I., Recht, B.: First-order methods almost always avoid strict saddle points · 2019
Later among the works it cites.
In: The 22nd International Conference on Artificial Intelligence and Statistics, pp. 2485–2494. PMLR (2019)
Mokhtari, A., Ozdaglar, A., Jadbabaie, A.: Efficient nonconvex empirical risk minimization via adaptive sample size methods · 2019
Later among the works it cites.
Mathematical Programming 176
O’Neill, M., Wright, S.J.: Behavior of accelerated gradient methods near critical points of nonconvex functions · 2019
Later among the works it cites.
arXiv preprint arXiv:1904.09237 (2019)
Reddi, S.J., Kale, S., Kumar, S.: On the convergence of adam and beyond · 2019
Later among the works it cites.
Zou, F., Shen, L., Jie, Z., Zhang, W., Liu, W.: A sufficient condition for convergences of adam and rmsprop · 2019
Later among the works it cites.
Mathematical Programming 180
Apidopoulos, V., Aujol, J.F., Dossal, C.: Convergence rate of inertial forward–backward algorithm beyond nesterov’s rule · 2020
Later among the works it cites.
arXiv preprint arXiv:2012.04061 (2020)
Das, R., Acharya, A., Hashemi, A., Sanghavi, S., Dhillon, I.S., Topcu, U.: Faster non-convex federated learning via global and local momentum · 2020
Later among the works it cites.
Advances in Neural Information Processing Systems 33
Gao, X., Gurbuzbalaban, M., Zhu, L.: Breaking reversibility accelerates langevin dynamics for non-convex optimization · 2020
Later among the works it cites.
Advances in Neural Information Processing Systems 33
Liu, Y., Gao, Y., Yin, W.: An improved analysis of stochastic gradient descent with momentum · 2020
Later among the works it cites.
Foundations of Computational Mathematics 20
Ma, C., Wang, K., Chi, Y., Chen, Y.: Implicit regularization in nonconvex statistical estimation: Gradient descent converges linearly for phase retrieval, matrix completion, and blind deconvolution · 2020
Later among the works it cites.
Springer Nature (2021)
Braun, P., Grüne, L., Kellett, C.M.: (In-) Stability of Differential Inclusions: Notions, Equivalences, and Lyapunov-like Characterizations · 2021
Later among the works it cites.
arXiv preprint arXiv:2108.11832 (2021)
Davis, D., Drusvyatskiy, D., Jiang, L.: Subgradient methods near active manifolds: saddle point avoidance, local convergence, and asymptotic normality · 2021
Later among the works it cites.
arXiv preprint arXiv:2102.09385 (2021)
Dereich, S., Kassing, S.: Convergence of stochastic gradient descent schemes for lojasiewicz-landscapes · 2021
Later among the works it cites.
Computational Mathematics and Mathematical Physics 61
Kurochkin, S.V.: Neural network with smooth activation functions and without bottlenecks is almost surely a morse function · 2021
Later among the works it cites.
IEEE Journal of Selected Topics in Signal Processing 15
Vial, P.H., Magron, P., Oberlin, T., Févotte, C.: Phase retrieval with bregman divergences and application to audio signal recovery · 2021
Later among the works it cites.
arXiv preprint arXiv:2106.02985 (2021)
Wang, J.K., Lin, C.H., Abernethy, J.: Escaping saddle points faster with stochastic momentum · 2021
Later among the works it cites.
Asymptotic Analysis 122
Yang, J., Hu, W., Li, C.J.: On the fast convergence of random perturbations of the gradient flow · 2021
Later among the works it cites.
Advances in Neural Information Processing Systems 34
Zhang, C., Li, T.: Escape saddle points by a simple gradient-descent based algorithm · 2021
Later among the works it cites.
arXiv preprint arXiv:2204.11292 (2022)
Can, B., Gurbuzbalaban, M.: Entropic risk-averse generalized momentum methods · 2022
Later among the works it cites.
IEEE Transactions on Information Theory pp. 1–1 (2022)
Dixit, R., Gürbüzbalaban, M., Bajwa, W.U.: Boundary conditions for linear exit time gradient trajectories around saddle points: Analysis and algorithm · 2022
Later among the works it cites.
Information and Inference: A Journal of the IMA (2022)
Dixit, R., Gürbüzbalaban, M., Bajwa, W.U.: Exit Time Analysis for Approximations of Gradient Descent Trajectories Around Saddle Points · 2022
Later among the works it cites.
Operations Research 70
Gao, X., Gürbüzbalaban, M., Zhu, L.: Global convergence of stochastic gradient hamiltonian monte carlo for nonconvex stochastic optimization: Nonasymptotic performance bounds and momentum-based acceleration · 2022
Later among the works it cites.
http://math.huji.ac.il/~mhochman/courses/fractals-2012/convergence-of-sets-and-measures.pdf
Hochman, M.: Convergence of sets and measures · 2022
Later among the works it cites.
http://math.huji.ac.il/~mhochman/courses/fractals-2012/
Hochman, M.: Convergence of sets and measures · 2022
Later among the works it cites.
IEEE Control Systems Magazine 42
Lessard, L.: The analysis of optimization algorithms: A dissipativity approach · 2022
Later among the works it cites.