Fetching the paper…
Reading the bibliography…
Adaptive gradient methods such as AdaGrad and its variants update the stepsize in stochastic gradient descent on the fly according to the gradients received along the way; such methods have gained widespread use in large-scale optimization for their ability to converge robustly, without the need to fine-tune the stepsize schedule.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A. Nemirovski and D. Yudin · 1983
Earlier work this paper cites.
Two-point step size gradient method
J. Barzilai and J. Borwein · 1988
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Introductory lectures on convex programming volume i: Basic course
Y. Nesterov · 1998
Earlier work this paper cites.
Numerical Optimization
S. Wright and J. Nocedal · 2006
Earlier work this paper cites.
Information-theoretic lower bounds on the oracle complexity of convex optimization
A. Agarwal, M. Wainwright, P. Bartlett, and P. Ravikumar · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
B. McMahan and M. Streeter · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Neural networks for machine learning-lecture 6a-overview of mini-batch gradient descent, 2012
G. Hinton N. Srivastava and K. Swersky · 2012
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method
M. Zeiler · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
S. Bubeck et al · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Cited alongside, same era.
Scale-free algorithms for online linear optimization
F. Orabona and D. Pal · 2015
Cited alongside, same era.
Improved svrg for non-strongly-convex or sum-of-non-convex objectives
Z. Allen-Zhu and Y. Yang · 2016
Cited alongside, same era.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
S. Ghadimi and G. Lan · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Non-convex finite-sum optimization via scsg methods
L. Lei, Cheng J., J. Chen, and M. Jordan · 2017
Later among the works it cites.
Online to offline conversions, universality and adaptive minibatch sizes
K. Levy · 2017
Later among the works it cites.
Variants of RMSProp and Adagrad with logarithmic regret bounds
M. C. Mukkamala and M. Hein · 2017
Later among the works it cites.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, and A. Antiga, L.and Lerer · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
A. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht · 2017
Later among the works it cites.
Natasha 2: Faster non-convex optimization than sgd
Z. Allen-Zhu · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fast incremental method for smooth nonconvex optimization
S. J. Reddi, S. Sra, B. Póczos, and A. Smola · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
T. Salimans and D. Kingma · 2016
Cited alongside, same era.
Barzilai-borwein step size for stochastic gradient descent
C. Tan, S. Ma, Y. Dai, and Y. Qian · 2016
Cited alongside, same era.
Finding approximate local minima faster than gradient descent
N. Agarwal, Z. Allen-Zhu, B. Bullins, and T. Hazan, E.and Ma · 2017
Cited alongside, same era.
Natasha: Faster non-convex stochastic optimization via strongly non-convex parameter
Z. Allen-Zhu · 2017
Cited alongside, same era.
“convex until proven guilty”: Dimension-free acceleration of gradient descent on non-convex functions
Y. Carmon, J. Duchi, O. Hinder, and A Sidford · 2017
Cited alongside, same era.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Closest in time.
Accelerated methods for nonconvex optimization
Y. Carmon, J. Duchi, O. Hinder, and A. Sidford · 2018
Closest in time.
Closing the generalization gap of adaptive gradient methods in training deep neural networks
J. Chen and Q. Gu · 2018
Closest in time.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
C. Fang, C. J. Li, Z. Lin, and T. Zhang · 2018
Closest in time.
On the convergence of adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Closest in time.
WNGrad: Learn the learning rate in gradient descent
X. Wu, R. Ward, and L. Bottou · 2018
Closest in time.
First-order stochastic algorithms for escaping from saddle points in almost linear time
Yi Xu, Rong Jin, and Tianbao Yang · 2018
Closest in time.
Stochastic nested variance reduced gradient descent for nonconvex optimization
D. Zhou, P. Xu, and Q. Gu · 2018
Closest in time.
Lower bounds for finding stationary points i
Y. Carmon, J. Duchi, O. Hinder, and A. Sidford · 2019
Closest in time.
On the convergence of stochastic gradient descent with adaptive stepsizes
X. Li and F. Orabona · 2019
Closest in time.