Fetching the paper…
Reading the bibliography…
The main result of this paper is: {\bf Theorem.} Let $f:\mathbb{R}^k\rightarrow \mathbb{R}$ be a $C^{1}$ function, so that $\nabla f$ is locally Lipschitz continuous.
H. B. Curry, The method of steepest descent for non-linear minimization problems
1944
Earlier work this paper cites.
H. Robbins and S. Monro, A stochastic approximation method
1951
Earlier work this paper cites.
J. B. Crockett and H. Chernoff, Gradient methods of maximization
1955
Earlier work this paper cites.
A. A. Goldstein, Cauchy’s method of minimization
1962
Earlier work this paper cites.
L. Armijo, Minimization of functions having Lipschitz continuous first partial derivatives
1966
Earlier work this paper cites.
P. Wolfe, Convergence conditions for ascent methods
1969
Earlier work this paper cites.
M. D. Asic and D. D. Adamovic, Limit points of sequences in metric spaces
1970
Earlier work this paper cites.
U. Helmke and J. B. Moore, Optimization and dynamical systems
1996
Earlier work this paper cites.
D. H. Wolpert and W. G. Macready, No free lunch theorems for optimisation, IEEE Transactions on evolutionary computation
1997
Cited alongside, same era.
D. P. Bertsekas, Nonlinear programming
1999
Cited alongside, same era.
J. Nocedal and S. J. Wright, Numerical optimization
1999
Cited alongside, same era.
A. Pinkus, Approximation theory of the MLP model in neural networks
1999
Cited alongside, same era.
Y. Nesterov, Introductory lectures on convex optimization : a basic course
2004
Cited alongside, same era.
P.-A. Absil, R. Mahony and B. Andrews, Convergence of the iterates of descent methods for analytic cost functions
2005
Cited alongside, same era.
L. N.Smith, No more pesky learning rate guessing games
2015
Later among the works it cites.
J. D. Lee, M. Simchowitz, M. I. Jordan and B. Recht, Gradient descent only converges to minimizers, JMRL: Workshop and conference proceedings, vol 49 (2016), 1–12
2016
Later among the works it cites.
M. Mahrsereci and P. Hennig, Probabilistic line searches for stochastic optimisation
2017
Later among the works it cites.
I. Panageas and G. Piliouras, Gradient descent only converges to minimizers: Non-isolated critical points and invariant regions
2017
Later among the works it cites.
N. Papernot, P. McDaniel, I. Goodfellow, S. Jha, Z. B. Celik and A. Swami, Practical black-box attacks against machine learning
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Boyd and L. Vandenberghe, Convex optimization
2009
Cited alongside, same era.
K. Lange, Optimization
2013
Cited alongside, same era.
L. Bottou, F. E. Curtis and J. Nocedal, Optimization methods for large-scale machine learning,
Cited in the paper.
A. J. Bray and and D. S. Dean, Statistics of critical points of gaussian fields on large-dimensional spaces
Cited in the paper.
A. Cauchy, Method général pour la résolution des systemes d’équations simulanées
Cited in the paper.
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli and Y. Bengjo, Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Cited in the paper.
K. Eyholt, I. Evtimov, E. Fernades, B. Li, A. Rahmati, C. Xiao, A. Prakash, T. Kohno and D. Song, Robust physical-world attacks on Deep Learning visual classification
2018
Later among the works it cites.
S. G. Finlayson, J. D. Bowers, J. Ito, J. L. Zittrain, A. L. Beam and I. S. Kohane, Adversarial attacks on medical machine learning
2019
Closest in time.