Fetching the paper…
Reading the bibliography…
Tuning hyperparameters of learning algorithms is hard because gradients are usually unavailable.
Backpropagation applied to handwritten zip code recognition
LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W., and Jackel, L. D · 1989
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Yoshua, Simard, Patrice, and Frasconi, Paolo · 1994
Earlier work this paper cites.
Automatic relevance determination for neural networks
MacKay, David J.C. and Neal, Radford M · 1994
Earlier work this paper cites.
Fast exact multiplication by the Hessian
Pearlmutter, Barak A · 1994
Earlier work this paper cites.
Building a better leapfrog
Hut, P., Makino, J., and McMillan, S · 1995
Earlier work this paper cites.
An investigation of the gradient descent process in neural networks
Pearlmutter, Barak · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Adaptive regularization in neural network modeling
Larsen, Jan, Svarer, Claus, Andersen, Lars Nonboe, and Hansen, Lars Kai · 1998
Earlier work this paper cites.
Optimal use of regularization and cross-validation in neural network modeling
Chen, Dingding and Hagan, Martin T · 1999
Earlier work this paper cites.
Gradient based adaptive regularization
Eigenmann, Robert and Nossek, Josef A · 1999
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Bengio, Yoshua · 2000
Earlier work this paper cites.
Choosing multiple parameters for support vector machines
Chapelle, Olivier, Vapnik, Vladimir, Bousquet, Olivier, and Mukherjee, Sayan · 2002
Earlier work this paper cites.
Information theory, inference, and learning algorithms
MacKay, David J.C · 2003
Earlier work this paper cites.
Derivative observations in Gaussian process models of dynamic systems
Solak, E., Murray Smith, R., Leithead, W.E., Leith, D., and Rasmussen, Carl E · 2003
Earlier work this paper cites.
Gaussian Processes for Machine Learning
Rasmussen, Carl E. and Williams, Christopher K.I · 2006
Cited alongside, same era.
Adaptive optimization of hyperparameters in L2-regularised logistic regression
Abdel-Gawad, Ahmed and Ratner, Simon · 2007
Cited alongside, same era.
Python for scientific computing
Oliphant, Travis E · 2007
Cited alongside, same era.
Efficient multiple hyperparameter learning for log-linear models
Foo, Chuan-sheng, Do, Chuong B., and Ng, Andrew Y · 2008
Cited alongside, same era.
Reverse-mode AD in a functional framework: Lambda the ultimate backpropagator
Pearlmutter, Barak A. and Siskind, Jeffrey Mark · 2008
Cited alongside, same era.
Theano: a CPU and GPU math expression compiler
Bergstra, James, Breuleux, Olivier, Bastien, Frédéric, Lamblin, Pascal, Pascanu, Razvan, Desjardins, Guillaume, Turian, Joseph, Warde-Farley, David, and Bengio, Yoshua · 2010
Practical Bayesian optimization of machine learning algorithms
Snoek, Jasper, Larochelle, Hugo, and Adams, Ryan P · 2012
Later among the works it cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Later among the works it cites.
Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures
Bergstra, James, Yamins, Daniel, and Cox, David · 2013
Later among the works it cites.
On the importance of initialization and momentum in deep learning
Sutskever, Ilya, Martens, James, Dahl, George, and Hinton, Geoffrey · 2013
Later among the works it cites.
Automatic differentiation of algorithms for machine learning
Baydin, A. G. and Pearlmutter, B. A · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Why does unsupervised pre-training help deep learning?
Erhan, Dumitru, Bengio, Yoshua, Courville, Aaron, Manzagol, Pierre-Antoine, Vincent, Pascal, and Bengio, Samy · 2010
Cited alongside, same era.
Algorithms for hyper-parameter optimization
Bergstra, James, Bardenet, Rémi, Bengio, Yoshua, Kégl, Balázs, et al · 2011
Cited alongside, same era.
Sequential model-based optimization for general algorithm configuration
Hutter, Frank, Hoos, Holger H, and Leyton-Brown, Kevin · 2011
Cited alongside, same era.
Theano: new features and speed improvements
Bastien, Frédéric, Lamblin, Pascal, Pascanu, Razvan, Bergstra, James, Goodfellow, Ian J., Bergeron, Arnaud, Bouchard, Nicolas, and Bengio, Yoshua · 2012
Cited alongside, same era.
Generic methods for optimization-based modeling
Domke, Justin · 2012
Cited alongside, same era.
Training deep and recurrent networks with hessian-free optimization
Martens, James and Sutskever, Ilya · 2012
Cited alongside, same era.
Courbariaux, Matthieu, Bengio, Yoshua, and David, Jean-Pierre · 2014
Later among the works it cites.
Multi-task neural networks for QSAR predictions
Dahl, George E, Jaitly, Navdeep, and Salakhutdinov, Ruslan · 2014
Later among the works it cites.
Graves, Alex, Wayne, Greg, and Danihelka, Ivo · 2014
Later among the works it cites.
Nested variational compression in deep Gaussian processes
Hensman, James and Lawrence, Neil D · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Later among the works it cites.
Towards more human-like concept learning in machines: Compositionality, causality, and learning-to-learn
Lake, Brenden M · 2014
Later among the works it cites.
Markov chain Monte Carlo and variational inference: Bridging the gap
Salimans, Tim, Kingma, Diederik P., and Welling, Max · 2014
Later among the works it cites.
Sequence to sequence learning with neural networks
Sutskever, Ilya, Vinyals, Oriol, and Le, Quoc V. V · 2014
Later among the works it cites.