Fetching the paper…
Reading the bibliography…
We develop and analyze a variant of Nesterov's accelerated gradient descent (AGD) for minimization of smooth non-convex functions.
Note sur la convergence de directions conjugées
E. Polak and G. Ribière · 1969
Earlier work this paper cites.
The fitting of power series, meaning polynomials, illustrated on band-spectroscopic data
A. E. Beaton and J. W. Tukey · 1974
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
A. Nemirovski and D. Yudin · 1983
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate O ( 1 / k 2 ) {O}(1/k^{2})
Y. Nesterov · 1983
Earlier work this paper cites.
Learning internal representations by error propagation
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Some NP-complete problems in quadratic and nonlinear programming
K. Murty and S. Kabadi · 1987
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
D. C. Liu and J. Nocedal · 1989
Earlier work this paper cites.
The MNIST database of handwritten digits, 1998
Y. LeCun, C. Cortes, and C. J. Burges · 1998
Earlier work this paper cites.
Optimization II: Standard numerical methods for nonlinear continuous optimization
A. Nemirovski · 1999
Earlier work this paper cites.
Squared functional systems and optimization problems
Y. Nesterov · 2000
Earlier work this paper cites.
Introductory Lectures on Convex Optimization
Y. Nesterov · 2004
Cited alongside, same era.
A survey of nonlinear conjugate gradient methods
W. W. Hager and H. Zhang · 2006
Cited alongside, same era.
Cubic regularization of Newton method and its global performance
Y. Nesterov and B. T. Polyak · 2006
Cited alongside, same era.
Numerical Optimization
J. Nocedal and S. J. Wright · 2006
Cited alongside, same era.
Supervised dictionary learning
J. Mairal, F. Bach, J. Ponce, G. Sapiro, and A. Zisserman · 2008
Cited alongside, same era.
Gradient-based algorithms with applications to signal recovery
A. Beck and M. Teboulle · 2009
Cited alongside, same era.
Matrix factorization techniques for recommender systems
Convex optimization: Algorithms and complexity
S. Bubeck · 2014
Later among the works it cites.
Phase retrieval via Wirtinger flow: Theory and algorithms
E. J. Candès, X. Li, and M. Soltanolkotabi · 2015
Later among the works it cites.
Qualitatively characterizing neural network optimization problems
I. J. Goodfellow, O. Vinyals, and A. M. Saxe · 2015
Later among the works it cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Later among the works it cites.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Later among the works it cites.
Adaptive restart for accelerated gradient schemes
B. O’Donoghue and E. Candès · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Koren, R. Bell, and C. Volinsky · 2009
Cited alongside, same era.
On the complexity of steepest descent, Newton’s and regularized Newton’s methods for nonconvex unconstrained optimization problems
C. Cartis, N. I. Gould, and P. L. Toint · 2010
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Cited alongside, same era.
How to make the gradients small
Y. Nesterov · 2012
Cited alongside, same era.
N. Agarwal, Z. Allen-Zhu, B. Bullins, E. Hazan, and T. Ma · 2016
Later among the works it cites.
Accelerated methods for non-convex optimization
Y. Carmon, J. C. Duchi, O. Hinder, and A. Sidford · 2016
Later among the works it cites.
Gradient descent only converges to minimizers
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht · 2016
Later among the works it cites.
Solving systems of random quadratic equations via truncated amplitude flow
G. Wang, G. B. Giannakis, and Y. C. Eldar · 2016
Later among the works it cites.