Fetching the paper…
Reading the bibliography…
We introduce a general method for improving the convergence rate of gradient-based optimizers that is easy to implement and works well in practice.
Increased rates of convergence through learning rate adaptation
R. A. Jacobs · 1988
Earlier work this paper cites.
Gain adaptation beats least squares?
R. S. Sutton · 1992
Earlier work this paper cites.
A direct adaptive method for faster backpropagation learning: The RPROP algorithm
M. Riedmiller and H. Braun · 1993
Earlier work this paper cites.
Fast exact multiplication by the Hessian
B. A. Pearlmutter · 1994
Earlier work this paper cites.
Parameter adaptation in stochastic optimization
L. B. Almeida, T. Langlois, J. D. Amaral, and A. Plakhov · 1998
Earlier work this paper cites.
An improved backpropagation method with adaptive learning rate
V. P. Plagianakos, D. G. Sotiropoulos, and M. N. Vrahatis · 1998
Earlier work this paper cites.
Local gain adaptation in stochastic gradient descent
N. N. Schraudolph · 1999
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Y. Bengio · 2000
Earlier work this paper cites.
Rates of convergence of adaptive step-size of stochastic approximation algorithms
S. Shao and P. P. C. Yip · 2000
Earlier work this paper cites.
Learning rate adaptation in stochastic gradient descent
V. P. Plagianakos, G. D. Magoulas, and M. N. Vrahatis · 2001
Earlier work this paper cites.
Fast online policy gradient learning with SMD gain vector adaptation
N. N. Schraudolph, J. Yu, and D. Aberdeen · 2006
Cited alongside, same era.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Cited alongside, same era.
Torch7: A MATLAB-like environment for machine learning
R. Collobert, K. Kavukcuoglu, and C. Farabet · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Practical recommendations for gradient-based training of deep architectures
Y. Bengio · 2012
Cited alongside, same era.
Random search for hyper-parameter optimization
J. Bergstra and Y. Bengio · 2012
Cited alongside, same era.
An evaluation of sequential model-based optimization for expensive blackbox functions
F. Hutter, H. Hoos, and K. Leyton-Brown · 2013
Later among the works it cites.
No more pesky learning rates
T. Schaul, S. Zhang, and Y. LeCun · 2013
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Later among the works it cites.
Gradient-based hyperparameter optimization through reversible learning
D. Maclaurin, D. K. Duvenaud, and R. P. Adams · 2015
Later among the works it cites.
Practical methodology
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generic methods for optimization-based modeling
J. Domke · 2012
Cited alongside, same era.
Practical Bayesian optimization of machine learning algorithms
J. Snoek, H. Larochelle, and R. P. Adams · 2012
Cited alongside, same era.
Lecture 6.5 – RMSProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures
J. Bergstra, D. Yamins, and D. D. Cox · 2013
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the Polyak-Lojasiewicz condition
H. Karimi, J. Nutini, and M. Schmidt · 2016
Later among the works it cites.
Convergence Analysis of an Adaptive Method of Gradient Descent
D. Martínez · 2017
Closest in time.
Automatic differentiation in PyTorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Closest in time.
Automatic differentiation in machine learning: a survey
A. G. Baydin, B. A. Pearlmutter, A. A. Radul, and J. M. Siskind · 2018
Closest in time.