Fetching the paper…
Reading the bibliography…
Most models in machine learning contain at least one hyperparameter to control for model complexity.
Some comments on C p C_{p}
Colin L Mallows · 1973
Earlier work this paper cites.
Estimation of the mean of a multivariate normal distribution
Charles M Stein · 1981
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Design and regularization of neural networks: the optimal use of a validation set
Jan Larsen, Lars Kai Hansen, Claus Svarer, and M Ohlsson · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
Adaptive regularization in neural network modeling
Jan Larsen, Claus Svarer, Lars Nonboe Andersen, and Lars Kai Hansen · 1998
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Yoshua Bengio · 2000
Earlier work this paper cites.
Choosing multiple parameters for support vector machines
Olivier Chapelle, Vladimir Vapnik, Olivier Bousquet, and Sayan Mukherjee · 2002
Earlier work this paper cites.
Accuracy and stability of numerical algorithms
Nicholas J Higham · 2002
Earlier work this paper cites.
Introductory lectures on convex optimization
Yurii Nesterov · 2004
Earlier work this paper cites.
Smooth optimization with approximate gradient
Alexandre d’Aspremont · 2008
Earlier work this paper cites.
Efficient multiple hyperparameter learning for log-linear models
Chuan-Sheng Foo, Chuong B. Do, and Andrew Y. Ng · 2008
Cited alongside, same era.
Trust region newton method for logistic regression
Chih-Jen Lin, Ruby C Weng, and S Sathiya Keerthi · 2008
Cited alongside, same era.
Cross-validation optimization for large scale structured classification kernel methods
Matthias W Seeger · 2008
Cited alongside, same era.
Gradient-based algorithms with applications to signal recovery
Amir Beck and Marc Teboulle · 2009
Cited alongside, same era.
Eric Brochu, Vlad M Cora, and Nando De Freitas · 2010
Cited alongside, same era.
Accurate telemonitoring of parkinson’s disease progression by noninvasive speech tests
Hybrid deterministic-stochastic methods for data fitting
Michael Friedlander and Mark Schmidt · 2012
Later among the works it cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Later among the works it cites.
A bilevel optimization approach for parameter learning in variational models
Karl Kunisch and Thomas Pock · 2013
Later among the works it cites.
Proximal algorithms
Neal Parikh and Stephen Boyd · 2013
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2013
Later among the works it cites.
Stein unbiased gradient estimator of the risk (SUGAR) for multiple parameter selection
Charles-Alban Deledalle, Samuel Vaiter, Jalal Fadili, and Gabriel Peyré · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Athanasios Tsanas, Max A Little, Patrick E McSharry, and Lorraine O Ramig · 2010
Cited alongside, same era.
Algorithms for hyper-parameter optimization
James S. Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl · 2011
Cited alongside, same era.
Parametric or nonparametric? a parametricness index for model selection
Wei Liu and Yuhong Yang · 2011
Cited alongside, same era.
Convergence rates of inexact proximal-gradient methods for convex optimization
Mark Schmidt, Nicolas Le Roux, and Francis R Bach · 2011
Cited alongside, same era.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio · 2012
Cited alongside, same era.
Generic methods for optimization-based modeling
Justin Domke · 2012
Cited alongside, same era.
Bilevel optimization with nonsmooth lower level problems
Peter Ochs, René Ranftl, Thomas Brox, and Thomas Pock
Cited in the paper.
Later among the works it cites.
Sequential model-based ensemble optimization
Alexandre Lacoste, Hugo Larochelle, François Laviolette, and Mario Marchand · 2014
Later among the works it cites.
Freeze-thaw bayesian optimization
Kevin Swersky, Jasper Snoek, and Ryan Prescott Adams · 2014
Later among the works it cites.
Bilevel approaches for learning of variational imaging models
Luca Calatroni, Cao Chung, Juan Carlos De Los Reyes, Carola-Bibiane Schönlieb, and Tuomo Valkonen · 2015
Later among the works it cites.
The structure of optimal parameters for image restoration problems
J.C. De los Reyes, Carola-Bibiane Schönlieb, and Tuomo Valkonen · 2015
Later among the works it cites.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan P. Adams · 2015
Later among the works it cites.