Fetching the paper…
Reading the bibliography…
We study two procedures (reverse-mode and forward-mode) for computing the gradient of the validation error with respect to the hyperparameters of any iterative learning algorithm such as stochastic gradient descent.
A Theoretical Framework for Back-Propagation
LeCun, Yann · 1988
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Williams, Ronald J. and Zipser, David · 1989
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Werbos, Paul J · 1990
Earlier work this paper cites.
DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST speech disc 1-1.1
Garofolo, John S., Lamel, Lori F., Fisher, William M., Fiscus, Jonathon G., and Pallett, David S · 1993
Earlier work this paper cites.
Gradient calculations for dynamic recurrent neural networks: A survey
Pearlmutter, Barak A · 1995
Earlier work this paper cites.
Design and regularization of neural networks: the optimal use of a validation set
Larsen, Jan, Hansen, Lars Kai, Svarer, Claus, and Ohlsson, M · 1996
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Bengio, Yoshua · 2000
Earlier work this paper cites.
Learning multiple tasks with kernel methods
Evgeniou, Theodoros, Micchelli, Charles A, and Pontil, Massimiliano · 2005
Earlier work this paper cites.
Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation
Griewank, Andreas and Walther, Andrea · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, Alex and Hinton, Geoffrey · 2009
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
Bergstra, James S., Bardenet, Rémi, Bengio, Yoshua, and Kégl, Balázs · 2011
Cited alongside, same era.
Learning output kernels with block coordinate descent
Dinuzzo, Francesco, Ong, Cheng S, Pillonetto, Gianluigi, and Gehler, Peter V · 2011
Cited alongside, same era.
Sequential model-based optimization for general algorithm configuration
Hutter, Frank, Hoos, Holger H., and Leyton-Brown, Kevin · 2011
Cited alongside, same era.
Learning with whom to share in multi-task feature learning
Kang, Zhuoliang, Grauman, Kristen, and Sha, Fei · 2011
Cited alongside, same era.
Random search for hyper-parameter optimization
Bergstra, James and Bengio, Yoshua · 2012
Cited alongside, same era.
Generic methods for optimization-based modeling
Domke, Justin · 2012
Cited alongside, same era.
Automatic differentiation in machine learning: a survey
Baydin, Atilim Gunes, Pearlmutter, Barak A., Radul, Alexey Andreyevich, and Siskind, Jeffrey Mark · 2015
Later among the works it cites.
Beyond Manual Tuning of Hyperparameters
Hutter, Frank, Lücke, Jörg, and Schmidt-Thieme, Lars · 2015
Later among the works it cites.
Efficient output kernel learning for multiple tasks
Jawanpuria, Pratik, Lapin, Maksim, Hein, Matthias, and Schiele, Bernt · 2015
Later among the works it cites.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, Dougal, Duvenaud, David K, and Adams, Ryan P · 2015
Later among the works it cites.
Scalable Bayesian Optimization Using Deep Neural Networks
Snoek, Jasper, Rippel, Oren, Swersky, Kevin, Kiros, Ryan, Satish, Nadathur, Sundaram, Narayanan, Patwary, Md Mostofa Ali, Prabhat, Mr, and Adams, Ryan P · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Practical bayesian optimization of machine learning algorithms
Snoek, Jasper, Larochelle, Hugo, and Adams, Ryan P · 2012
Cited alongside, same era.
Making a Science of Model Search: Hyperparameter Optimization in Hundreds of Dimensions for Vision Architectures
Bergstra, James, Yamins, Daniel, and Cox, David D · 2013
Cited alongside, same era.
Multi-task bayesian optimization
Swersky, Kevin, Snoek, Jasper, and Adams, Ryan P · 2013
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Kingma, Diederik and Ba, Jimmy · 2014
Cited alongside, same era.
Phonetic context embeddings for dnn-hmm phone recognition
Badino, Leonardo · 2016
Later among the works it cites.
Hyperparameter optimization with approximate gradient
Pedregosa, Fabian · 2016
Later among the works it cites.
Neural architecture search with reinforcement learning
Zoph, Barret and Le, Quoc V · 2016
Later among the works it cites.
Personal communication, 2017
Badino, Leonardo · 2017
Closest in time.
Fan, Yang, Tian, Fei, Qin, Tao, Bian, Jiang, and Liu, Tie-Yan · 2017
Closest in time.