Fetching the paper…
Reading the bibliography…
Working with any gradient-based machine learning algorithm involves the tedious task of tuning the optimizer's hyperparameters, such as its step size.
Energy and policy considerations for deep learning in NLP
E. Strubell, A. Ganesh, and A. McCallum · 1906
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Parameter adaptation in stochastic optimization
L. E. Almeida, T. Langlois, J. F. M. do Amaral, and A. Plakhov · 1999
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Y. Bengio · 2000
Earlier work this paper cites.
Caltech-256 object category dataset
G. Griffin, A. Holub, and P. Perona · 2007
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Generic methods for optimization-based modeling
J. Domke · 2012
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2012
Earlier work this paper cites.
ADADELTA: An adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Visualizing and understanding recurrent networks
A. Karpathy, J. Johnson, and L. Fei-Fei · 2015
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
D. Maclaurin, D. Duvenaud, and R. P. Adams · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Scalable gradient-based tuning of continuous regularization hyperparameters
J. Luketina, M. Berglund, K. Greff, and T. Raiko · 2016
Cited alongside, same era.
Hyperparameter optimization with approximate gradient
F. Pedregosa · 2016
Training deep networks without learning rates through coin betting
F. Orabona and T. Tommasi · 2017
Later among the works it cites.
Automatic differentiation in PyTorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Later among the works it cites.
Convergence analysis of an adaptive method of gradient descent
D. M. Rubio · 2017
Later among the works it cites.
Online learning rate adaptation with hypergradient descent
A. G. Baydin, R. Cornish, D. M. Rubio, M. Schmidt, and F. Wood · 2018
Later among the works it cites.
Deepobs: A deep learning optimizer benchmark suite
F. Schneider, L. Balles, and P. Hennig · 2018
Later among the works it cites.
Hyperparameter Optimization , pages 3–33
M. Feurer and F. Hutter · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Forward and reverse gradient-based hyperparameter optimization
L. Franceschi, M. Donini, P. Frasconi, and M. Pontil · 2017
Cited alongside, same era.
torch-rnn
J. Johnson · 2017
Cited alongside, same era.
Painless stochastic gradient: Interpolation, line-search, and convergence rates
S. Vaswani, A. Mishkin, I. Laradji, M. Schmidt, G. Gidel, and S. Lacoste-Julien · 2019
Closest in time.
Descending through a crowded valley-benchmarking deep learning optimizers
R. M. Schmidt, F. Schneider, and P. Hennig · 2021
Closest in time.