Fetching the paper…
Reading the bibliography…
We present a strikingly simple proof that two rules are sufficient to automate gradient descent: 1) don't increase the stepsize too fast and 2) don't overstep the local curvature.
“Cauchy’s method of minimization”
AA Goldstein · 1962
Earlier work this paper cites.
“An application of the method of gradient descent to the solution of the network transportation problem”
NZ Shor · 1962
Earlier work this paper cites.
“Gradient methods for minimizing functionals”
Boris Polyak · 1963
Earlier work this paper cites.
“Minimization of functions having Lipschitz continuous first partial derivatives.”
Larry Armijo · 1966
Earlier work this paper cites.
“Minimization of nonsmooth functionals”
Boris Polyak · 1969
Earlier work this paper cites.
“Problem complexity and method efficiency in optimization”
A.. Nemirovsky and D.. Yudin · 1983
Earlier work this paper cites.
“A method for unconstrained convex minimization problem with the rate of convergence O ( 1 / k 2 ) O(1/k^{2}) ”
Yurii Nesterov · 1983
Earlier work this paper cites.
“Two-point step size gradient methods”
Jonathan Barzilai and Jonathan. Borwein · 1988
Earlier work this paper cites.
“On the Barzilai and Borwein choice of steplength for the gradient method”
Marcos Raydan · 1993
Earlier work this paper cites.
“R-linear convergence of the Barzilai and Borwein gradient method”
Yu-Hong Dai and Li-Zhi Liao · 2002
Earlier work this paper cites.
“Cubic regularization of Newton method and its global performance”
Yurii Nesterov and Boris. Polyak · 2006
Earlier work this paper cites.
“Learning multiple layers of features from tiny images”, 2009
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
“Adaptive Bound Optimization for Online Convex Optimization”
H. McMahan and Matthew Streeter · 2010
Earlier work this paper cites.
“Distributed algorithms via gradient descent for Fisher markets”
Benjamin Birnbaum, Nikhil Devanur and Lin Xiao · 2011
Earlier work this paper cites.
“Adaptive subgradient methods for online learning and stochastic optimization”
John Duchi, Elad Hazan and Yoram Singer · 2011
Cited alongside, same era.
“Cauchy and the gradient method”
Claude Lemar\’echal · 2012
Cited alongside, same era.
“Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude”, COURSERA: Neural Networks for Machine Learning, 2012
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
“ADADELTA: an adaptive learning rate method”
Matthew. Zeiler · 2012
Cited alongside, same era.
“Convergence of descent methods for semi-algebraic and tame problems: proximal algorithms, forward–backward splitting, and regularized Gauss–Seidel methods”
Hedy Attouch, J\’er\ˆome Bolte and Benar Svaiter · 2013
Cited alongside, same era.
“The variable metric forward-backward splitting algorithm under mild differentiability assumptions”
Saverio Salzo · 2017
Later among the works it cites.
“Smooth strongly convex interpolation and exact worst-case performance of first-order methods”
Adrien. Taylor, Julien. Hendrickx and Francois Glineur · 2017
Later among the works it cites.
“signSGD: Compressed Optimisation for Non-Convex Problems”
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli and Animashree Anandkumar · 2018
Later among the works it cites.
“On the Convergence of Adam and Beyond”
Sashank. Reddi, Satyen Kale and Sanjiv Kumar · 2018
Later among the works it cites.
“Stabilized Barzilai-Borwein Method”
Oleg Burdakov, Yuhong Dai and Na Huang · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Gradient methods for minimizing composite functions”
Yu. Nesterov · 2013
Cited alongside, same era.
“Introductory lectures on convex optimization: A basic course”
Yurii Nesterov · 2013
Cited alongside, same era.
“Performance of first-order methods for smooth convex minimization: a novel approach”
Yoel Drori and Marc Teboulle · 2014
Cited alongside, same era.
“Adam: A Method for Stochastic Optimization”
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
“A descent lemma beyond Lipschitz gradient continuity: first-order methods revisited and applications”
Heinz Bauschke, J\’er\ˆome Bolte and Marc Teboulle · 2016
Cited alongside, same era.
“On the convergence of the forward–backward splitting method with linesearches”
Jos\’e Bello and Tran.A. Nghia · 2016
Cited alongside, same era.
“The movielens datasets: History and context”
F. Harper and Joseph. Konstan · 2016
Cited alongside, same era.
Elad Hazan and Sham Kakade · 2019
Closest in time.
“Error Feedback Fixes SignSGD and other Gradient Compression Schemes”
Sai Karimireddy, Quentin Rebjock, Sebastian Stich and Martin Jaggi · 2019
Closest in time.
“Dual space preconditioning for gradient descent”
Chris. Maddison, Daniel Paulin, Yee Teh and Arnaud Doucet · 2019
Closest in time.
“Golden ratio algorithms for variational inequalities”
Yura Malitsky · 2019
Closest in time.
“Unified Optimal Analysis of the (Stochastic) Gradient Method”
Sebastian. Stich · 2019
Closest in time.
“Stochastic first-order methods: non-asymptotic and computer-aided analyses via potential functions”
Adrien Taylor and Francis Bach · 2019
Closest in time.
“AdaGrad Stepsizes: Sharp Convergence Over Nonconvex Landscapes”
Rachel Ward, Xiaoxia Wu and Leon Bottou · 2019
Closest in time.
“Adaptive restart of accelerated gradient methods under local quadratic growth condition”
Olivier Fercoq and Zheng Qu · 2095
Closest in time.