Fetching the paper…
Reading the bibliography…
Finding the optimal hyperparameters of a model can be cast as a bilevel optimization problem, typically solved using zero-order techniques.
Methods of conjugate gradients for solving linear systems , volume 49
M. R. Hestenes and E. Stiefel · 1952
Earlier work this paper cites.
The convergence of the random search method in the extremal control of a many parameter system
L. A. Rastrigin · 1963
Earlier work this paper cites.
A simple automatic derivative evaluation program
R. E. Wengert · 1964
Earlier work this paper cites.
Estimating WAIS IQ from Shipley Scale scores: Another cross-validation
L. R. A. Stone and J.C. Ramer · 1965
Earlier work this paper cites.
Ridge regression: Biased estimation for nonorthogonal problems
A. E. Hoerl and R. W. Kennard · 1970
Earlier work this paper cites.
The representation of the cumulative rounding error of an algorithm as a Taylor expansion of the local rounding errors
S. Linnainmaa · 1970
Earlier work this paper cites.
A new look at the statistical model identification
H. Akaike · 1974
Earlier work this paper cites.
Estimating the dimension of a model
G. Schwarz · 1978
Earlier work this paper cites.
Distribution-free performance bounds for potential function rules
L. Devroye and T. Wagner · 1979
Earlier work this paper cites.
Splitting algorithms for the sum of two nonlinear operators
P-L. Lions and B. Mercier · 1979
Earlier work this paper cites.
Estimation of the mean of a multivariate normal distribution
C. M. Stein · 1981
Earlier work this paper cites.
How biased is the apparent error rate of a prediction rule?
B. Efron · 1986
Earlier work this paper cites.
Introduction to optimization
B. T. Polyak · 1987
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
D. C. Liu and J. Nocedal · 1989
Earlier work this paper cites.
The bayesian approach to local optimization
J. Mockus · 1989
Earlier work this paper cites.
A training algorithm for optimal margin classifiers
B. E. Boser, I. M. Guyon, and V. N. Vapnik · 1992
Earlier work this paper cites.
Measure theory and fine properties of functions
L. C. Evans and R. F. Gariepy · 1992
Earlier work this paper cites.
Automatic parameter selection by minimizing estimated error
R. Kohavi and G. H. John · 1995
Earlier work this paper cites.
Design and regularization of neural networks: the optimal use of a validation set
J. Larsen, L. K. Hansen, C. Svarer, and M. Ohlsson · 1996
Earlier work this paper cites.
Generalized hessian properties of regularized nonsmooth functions
R. A. Poliquin and R. T. Rockafellar · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
R. Tibshirani · 1996
Earlier work this paper cites.
Atomic decomposition by basis pursuit
S. S. Chen, D. L. Donoho, and M. A. Saunders · 1998
Earlier work this paper cites.
Efficient global optimization of expensive black-box functions
D. R. Jones, M. Schonlau, and W. J. Welch · 1998
Earlier work this paper cites.
Efficient backprop
Y. A. LeCun, L. Bottou, G. B. Orr, and K-R. Müller · 1998
Earlier work this paper cites.
Properties of support vector machines
M. Pontil and A. Verri · 1998
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
J. C. Platt · 1999
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Y. Bengio · 2000
Earlier work this paper cites.
Choosing multiple parameters for support vector machines
O. Chapelle, V. Vapnik, O. Bousquet, and S. Mukherjee · 2002
Earlier work this paper cites.
Accuracy and stability of numerical algorithms
N. J. Higham · 2002
Earlier work this paper cites.
Introductory lectures on convex optimization , volume 87 of Applied Optimization
Y. Nesterov · 2004
Earlier work this paper cites.
Signal recovery by proximal forward-backward splitting
P. L. Combettes and V. R. Wajs · 2005
Earlier work this paper cites.
Regularization and variable selection via the elastic net
H. Zou and T. J. Hastie · 2005
Earlier work this paper cites.
A mathematical model for automatic differentiation in machine learning
J. Bolte and E. Pauwels · 2006
Earlier work this paper cites.
Numerical optimization
J. Nocedal and S. J. Wright · 2006
Earlier work this paper cites.
An overview of bilevel optimization
B. Colson, P. Marcotte, and G. Savard · 2007
Earlier work this paper cites.
Identifying active manifolds
W. L. Hare and A. S. Lewis · 2007
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
J. D. Hunter · 2007
Earlier work this paper cites.
An interior-point method for large-scale l1-regularized logistic regression
K. Koh, S.-J. Kim, and S. Boyd · 2007
Earlier work this paper cites.
Sure independence screening for ultrahigh dimensional feature space
J. Fan and J. Lv · 2008
Cited alongside, same era.
Efficient multiple hyperparameter learning for log-linear models
C. S. Foo, C. B. Do, and A. Y. Ng · 2008
Cited alongside, same era.
Engineering design via surrogate modelling: a practical guide
A. Forrester, A. Sobester, and A. Keane · 2008
Cited alongside, same era.
Sup-norm convergence rate and sign concentration property of Lasso and Dantzig estimators
K. Lounici · 2008
Cited alongside, same era.
Cross-validation optimization for large scale structured classification kernel methods
M. W. Seeger · 2008
Cited alongside, same era.
Simultaneous analysis of Lasso and Dantzig selector
P. J. Bickel, Y. Ritov, and A. B. Tsybakov · 2009
Cited alongside, same era.
Conic optimization via operator splitting and homogeneous self-dual embedding
B. O’Donoghue, E. Chu, N. Parikh, and S. Boyd · 2016
Later among the works it cites.
Hyperparameter optimization with approximate gradient
F. Pedregosa · 2016
Later among the works it cites.
Optnet: Differentiable optimization as a layer in neural networks
B. Amos and J. Z. Kolter · 2017
Later among the works it cites.
Forward and reverse gradient-based hyperparameter optimization
L. Franceschi, M. Donini, P. Frasconi, and M. Pontil · 2017
Later among the works it cites.
Activity identification and local linear convergence of Forward–Backward-type Methods
J. Liang, J. Fadili, and G. Peyré · 2017
Later among the works it cites.
Breaking the nonsmooth barrier: A scalable parallel method for composite optimization
F. Pedregosa, R. Leblond, and S. Lacoste-Julien · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Block-coordinate gradient descent method for linearly constrained nonsmooth separable optimization
P. Tseng and S. Yun · 2009
Cited alongside, same era.
A survey of cross-validation procedures for model selection
S. Arlot and A. Celisse · 2010
Cited alongside, same era.
A tutorial on Bayesian optimization of expensive cost functions, with application to active user modeling and hierarchical reinforcement learning
E. Brochu, V. M. Cora, and N. De Freitas · 2010
Cited alongside, same era.
Regularization paths for generalized linear models via coordinate descent
J. Friedman, T. J. Hastie, and R. Tibshirani · 2010
Cited alongside, same era.
Square-root Lasso: pivotal recovery of sparse signals via conic programming
A. Belloni, V. Chernozhukov, and L. Wang · 2011
Cited alongside, same era.
Algorithms for hyper-parameter optimization
J. Bergstra, R. Bardenet, Y. Bengio, and B. Kégl · 2011
Cited alongside, same era.
Later among the works it cites.
Automatic differentiation in machine learning: a survey
A. G. Baydin, B. A. Pearlmutter, A. A. Radul, and J. M. Siskind · 2018
Later among the works it cites.
Sensitivity analysis for mirror-stratifiable convex functions
J. Fadili, J. Malick, and G. Peyré · 2018
Later among the works it cites.
Bilevel programming for hyperparameter optimization and meta-learning
L. Franceschi, P. Frasconi, S. Salzo, and M. Pontil · 2018
Later among the works it cites.
A tutorial on Bayesian optimization
P.I. Frazier · 2018
Later among the works it cites.
Bilevel learning of the group lasso structure
J. Frecon, S. Salzo, and M. Pontil · 2018
Later among the works it cites.
Approximation methods for bilevel programming
S. Ghadimi and M. Wang · 2018
Later among the works it cites.
Celer: a fast solver for the lasso with dual extrapolation
M. Massias, A. Gramfort, and J. Salmon · 2018
Later among the works it cites.
Greed is good: greedy optimization methods for large-scale structured problems
J. Nutini · 2018
Later among the works it cites.
Model consistency of partly smooth regularizers
S. Vaiter, G. Peyré, and J. Fadili · 2018
Later among the works it cites.
Differentiable convex optimization layers
A. Agrawal, B. Amos, S. Barratt, S. Boyd, S. Diamond, and J. Z. Kolter · 2019
Later among the works it cites.
Deep equilibrium models
S. Bai, J. Z. Kolter, and V. Koltun · 2019
Later among the works it cites.
Hyperparameter optimization
M. Feurer and F. Hutter · 2019
Later among the works it cites.
Optimizing millions of hyperparameters by implicit differentiation
J. Lorraine, P. Vicol, and D. Duvenaud · 2019
Later among the works it cites.
“active-set complexity” of proximal gradient: How long does it take to find the sparsity pattern?
J. Nutini, M. Schmidt, and W. Hare · 2019
Later among the works it cites.
SCS: Splitting conic solver, version 2.1.2, 2019
B. O’Donoghue, E. Chu, N. Parikh, and S. Boyd · 2019
Later among the works it cites.
Meta-learning with implicit gradients
A. Rajeswaran, C. Finn, S. M. Kakade, and S. Levine · 2019
Later among the works it cites.
Super-efficiency of automatic differentiation for functions defined as a minimum
P. Ablin, G. Peyré, and T. Moreau · 2020
Later among the works it cites.
Multiscale deep equilibrium models
S. Bai, V. Koltun, and J. Z. Kolter · 2020
Later among the works it cites.
Implicit differentiation of Lasso-type models for hyperparameter optimization
Q. Bertrand, Q. Klopfenstein, M. Blondel, S. Vaiter, A. Gramfort, and J. Salmon · 2020
Later among the works it cites.
Learning to solve TV regularised problems with unrolled algorithms
H. Cherkaoui, J. Sulam, and T. Moreau · 2020
Later among the works it cites.
On the iteration complexity of hypergradient computation
R. Grazzi, L. Franceschi, M. Pontil, and S. Salzo · 2020
Later among the works it cites.
Array programming with NumPy
C. R. Harris, K. J. Millman, S. J. van der Walt, R. Gommers, P. Virtanen, D. Cournapeau, E. Wieser, J. Taylor, S. Berg, N. J. Smith, R. Kern, M. Picus, S. Hoyer, M. H. van Kerkwijk, M. Brett, A. Haldane, J. Fern’andez del R’ıo, M. Wiebe, P. Peterson, P. G’erard-Marchant, K. Sheppard, T. Reddy, W. Weckesser, H. Abbasi, C. Gohlke, and T. E. Oliphant · 2020
Later among the works it cites.
Provably faster algorithms for bilevel optimization and applications to meta-learning
K. Ji, J. Yang, and Y. Liang · 2020
Later among the works it cites.
Model identification and local linear convergence of coordinate descent
Q. Klopfenstein, Q. Bertrand, A. Gramfort, J. Salmon, and S. Vaiter · 2020
Later among the works it cites.
A generic first-order algorithmic framework for bi-level programming beyond lower-level singleton
R. Liu, P. Mu, X. Yuan, S. Zeng, and J. Zhang · 2020
Later among the works it cites.
Dual extrapolation for sparse generalized linear models
M. Massias, S. Vaiter, A. Gramfort, and J. Salmon · 2020
Later among the works it cites.
Scipy 1.0: fundamental algorithms for scientific computing in python
P. Virtanen, R. Gommers, T. E. Oliphant, M. Haberland, T. Reddy, D. Cournapeau, E. Burovski, P. Peterson, W. Weckesser, J. Bright, et al · 2020
Later among the works it cites.
Machine learning refined: Foundations, algorithms, and applications
J. Watt, R. Borhani, and A. K. Katsaggelos · 2020
Later among the works it cites.
Neural monotone operator equilibrium networks
E. Winston and Z. Kolter · 2020
Later among the works it cites.
Efficient and modular implicit differentiation
M. Blondel, Q. Berthet, M. Cuturi, R. Frostig, S. Hoyer, F. Llinares-López, F. Pedregosa, and J.-P. Vert · 2021
Closest in time.
Convergence properties of stochastic hypergradients
R. Grazzi, M. Pontil, and S. Salzo · 2021
Closest in time.