Fetching the paper…
Reading the bibliography…
While Standard gradient descent is one very popular optimisation method, its convergence cannot be proven beyond the class of functions whose gradient is globally Lipschitz continuous.
H. B. Curry, The method of steepest descent for non-linear minimization problems, Quarterly of applied mathematics, 2 (October 1944), no 3, 258–261
1944
Earlier work this paper cites.
H. Robbins and S. Monro, A stochastic approximation method, Annals of Mathematical Statistics, vol 22, pp. 400–407, 1951
1951
Earlier work this paper cites.
J. B. Crockett and H. Chernoff, Gradient methods of maximization, Pacific J. Math. 5 (1955), 33–50
1955
Earlier work this paper cites.
A. A. Goldstein, Cauchy’s method of minimization, Numerische Mathematik 4 (1962), 146–150
1962
Earlier work this paper cites.
L. Armijo, Minimization of functions having Lipschitz continuous first partial derivatives, Pacific J. Math. 16 (1966), no. 1, 1–3
1966
Earlier work this paper cites.
H. W. Swann, A survey of non-linear optimisation techniques, FEBS letters, Volume 2, supplement 1, March 1969, pages S39–S55
1969
Earlier work this paper cites.
P. Wolfe, Convergence conditions for ascent methods, SIAM Review 11 (April 1969), no 2, 226–235
1969
Earlier work this paper cites.
M. D. Asic and D. D. Adamovic, Limit points of sequences in metric spaces, The American mathematical monthly, vol 77, so 6 (June–July 1970), 613–616
1970
Earlier work this paper cites.
U. Helmke and J. B. Moore, Optimization and dynamical systems
1996
Earlier work this paper cites.
D. H. Wolpert and W. G. Macready, No free lunch theorems for optimisation, IEEE Transactions on evolutionary computation, Vol. 1, No 1, April 1997, 67–82
1997
Earlier work this paper cites.
D. P. Bertsekas, Nonlinear programming, 2nd edition, Athena Scientific, Belmont, Massachusetts, 1999
1999
Earlier work this paper cites.
J. Nocedal and S. J. Wright, Numerical optimization, Springer series in operations research, 1999
1999
Earlier work this paper cites.
N. Qian, On the momentum term in gradient descent learning algorithms, Neural networks: the official journal of the International Neural Network Society, 12(1):145–151, 1999
1999
Earlier work this paper cites.
Y. Nesterov, Introductory lectures on convex optimization : a basic course, 2004, Kluwer Academic Publishers. ISBN 978-1402075537
2004
Earlier work this paper cites.
P.-A. Absil, R. Mahony and B. Andrews, Convergence of the iterates of descent methods for analytic cost functions, SIAM J. Optim. 16 (2005), vol 16, no 2, 531–547
2005
Cited alongside, same era.
S. Boyd and L. Vandenberghe, Convex optimization, 7th printing with corrections, Cambridge University Press, 2009
2009
Cited alongside, same era.
J. Duchi, E. Hazan, and Y. Singer, Adaptive subgradient methods for online learning and stochastic optimization, Journal of Machine Learning Research, 12:2121–2159, 2011
2011
Cited alongside, same era.
G. Hinton, N Srivastava, and K. Swersky, Lecture 6a overview of mini-batch gradient descent, Coursera Lecture slides https://class.coursera.org/neuralnets-2012-001/lecture, 2012
2012
Cited alongside, same era.
K. Lange, Optimization, 2nd edition, Springer texts in statistics, New York 2013
2013
2016
Later among the works it cites.
J. D. Lee, M. Simchowitz, M. I. Jordan and B. Recht, Gradient descent only converges to minimizers, JMRL: Workshop and conference proceedings, vol 49 (2016), 1–12
2016
Later among the works it cites.
K. Kawaguchi, Deep learning without local minima, part of Advances in Neural information processing system 29 (NIPS, 2016)
2016
Later among the works it cites.
J. Hu, L. Shen, and G. Sun. Squeeze-and-excitation networks. CoRR, abs/1709.01507, 2017
2017
Later among the works it cites.
S. Ioffe, Batch renormalization: towards reducing minibatch dependence in batch-normalized models, Nips, 2017
2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2014
Cited alongside, same era.
S. Shalev-Shwartz and S. Ben-David, Understanding machine learning: from theory to algorithms
2014
Cited alongside, same era.
R. Ge, F. Huang, C. Jin and Y. Yuan, Escaping from saddle points - online stochastic gradient for tensor decomposition, JMLR: Workshop and conference proceedings, vol 40: 1–46, 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
S. Ioffe and C. Szegedy, Batch normalization: accelerating deep network training by reducing internal covariate shift, Proceedings of the 32Nd International Conference on International Conference on Machine Learning - Volume 37, JMLR.org, 2015 , 448-456
2015
Cited alongside, same era.
D. P. Kingma and J. Lei Ba. Adam: a method for stochastic optimization, International Conference on Learning Representations, pages 1–13, 2015
2015
Cited alongside, same era.
M. A. Nielsen, Neural networks and deep learning, Determination Press, 2015
2015
Cited alongside, same era.
Later among the works it cites.
2017
Later among the works it cites.
M. Mahrsereci and P. Hennig, Probabilistic line searches for stochastic optimisation, JMLR, vol 18 (2017), 1–59
2017
Later among the works it cites.
I. Panageas and G. Piliouras, Gradient descent only converges to minimizers: Non-isolated critical points and invariant regions, 8th Innovations in theoretical computer science conference (ITCS 2017), Editor: C. H. Papadimitrou, article no 2, pp. 2:1–2:12, Leibniz international proceedings in informatics (LIPICS), Dagstuhl Publishing. Germany
2017
Later among the works it cites.
N. Papernot, P. McDaniel, I. Goodfellow, S. Jha, Z. B. Celik and A. Swami, Practical black-box attacks against machine learning
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
K. Eyholt, I. Evtimov, E. Fernades, B. Li, A. Rahmati, C. Xiao, A. Prakash, T. Kohno and D. Song, Robust physical-world attacks on Deep Learning visual classification
2018
Closest in time.