Fetching the paper…
Reading the bibliography…
We describe a general framework for online adaptation of optimization hyperparameters by `hot swapping' their values during learning.
Learning dynamic algorithm portfolios
Gagliolo, M. and Schmidhuber, J. (2006) · 2006
Earlier work this paper cites.
On upper-confidence bound policies for non-stationary bandit problems
Garivier, A. and Moulines, E. (2008) · 2008
Earlier work this paper cites.
On optimization methods for deep learning
Ngiam, J., Coates, A., Lahiri, A., Prochnow, B., Le, Q. V., and Ng, A. Y. (2011) · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y. (2012) · 2012
Cited alongside, same era.
A stochastic gradient method with an exponential convergence _rate for finite training sets
Roux, N. L., Schmidt, M., and Bach, F. R. (2012) · 2012
Cited alongside, same era.
Schaul, T., Zhang, S., and LeCun, Y. (2012) · 2012
Cited alongside, same era.
Practical bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P. (2012) · 2012
Later among the works it cites.
ADADELTA: An adaptive learning rate method
Zeiler, M. D. (2012) · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…