Fetching the paper…
Reading the bibliography…
Hyperparameter optimization of neural networks can be elegantly formulated as a bilevel optimization problem.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
J. Schmidhuber · 1992
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
M. Marcus, B. Santorini, and M. A. Marcinkiewicz · 1993
Earlier work this paper cites.
Design and regularization of neural networks: the optimal use of a validation set
J. Larsen, L. K. Hansen, C. Svarer, and M. Ohlsson · 1996
Earlier work this paper cites.
Introductory lectures on convex programming volume i: Basic course
Y. Nesterov · 1998
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Y. Bengio · 2000
Earlier work this paper cites.
Numerical optimization
J. Nocedal and S. Wright · 2006
Earlier work this paper cites.
An overview of bilevel optimization
B. Colson, P. Marcotte, and G. Savard · 2007
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky et al · 2009
Earlier work this paper cites.
Sequential model-based optimization for general algorithm configuration (extended version)
F. Hutter, H. H. Hoos, and K. Leyton-Brown · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
J. S. Bergstra, R. Bardenet, Y. Bengio, and B. Kégl · 2011
Earlier work this paper cites.
Deep sparse rectifier neural networks
X. Glorot, A. Bordes, and Y. Bengio · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
J. Bergstra and Y. Bengio · 2012
Earlier work this paper cites.
Generic methods for optimization-based modeling
J. Domke · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Deep boltzmann machines and the centering trick
G. Montavon and K.-R. Müller · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
J. Snoek, H. Larochelle, and R. P. Adams · 2012
Earlier work this paper cites.
J. Staines and D. Barber · 2012
Earlier work this paper cites.
Regularization of neural networks using dropconnect
L. Wan, M. Zeiler, S. Zhang, Y. Le Cun, and R. Fergus · 2013
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
New insights and perspectives on the natural gradient method
J. Martens · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Comparison of regularization methods for imagenet classification with deep convolutional neural networks
E. A. Smirnov, D. M. Timoshenko, and S. N. Andrianov · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Freeze-thaw bayesian optimization
K. Swersky, J. Snoek, and R. P. Adams · 2014
Cited alongside, same era.
UCI machine learning repository, 2017
D. Dua and C. Graff · 2017
Later among the works it cites.
Forward and reverse gradient-based hyperparameter optimization
L. Franceschi, M. Donini, P. Frasconi, and M. Pontil · 2017
Later among the works it cites.
Regularizing and optimizing lstm language models
S. Merity, N. S. Keskar, and R. Socher · 2017
Later among the works it cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017
H. Xiao, K. Rasul, and R. Vollgraf · 2017
Later among the works it cites.
Noisy natural gradient as variational inference
G. Zhang, S. Sun, D. Duvenaud, and R. Grosse · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
D. Maclaurin, D. Duvenaud, and R. Adams · 2015
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
J. Martens and R. Grosse · 2015
Cited alongside, same era.
Scalable bayesian optimization using deep neural networks
J. Snoek, O. Rippel, K. Swersky, R. Kiros, N. Satish, N. Sundaram, M. Patwary, M. Prabhat, and R. Adams · 2015
Cited alongside, same era.
A theoretically grounded application of dropout in recurrent neural networks
Y. Gal and Z. Ghahramani · 2016
Cited alongside, same era.
D. Ha, A. Dai, and Q. V. Le · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, and S. Wanderman-Milne · 2018
Later among the works it cites.
Bilevel programming for hyperparameter optimization and meta-learning
L. Franceschi, P. Frasconi, S. Salzo, R. Grazzi, and M. Pontil · 2018
Later among the works it cites.
Massively parallel hyperparameter tuning
L. Li, K. Jamieson, A. Rostamizadeh, E. Gonina, M. Hardt, B. Recht, and A. Talwalkar · 2018
Later among the works it cites.
Tune: A research platform for distributed model selection and training
R. Liaw, E. Liang, R. Nishihara, P. Moritz, J. E. Gonzalez, and I. Stoica · 2018
Later among the works it cites.
Darts: Differentiable architecture search
H. Liu, K. Simonyan, and Y. Yang · 2018
Later among the works it cites.
Stochastic hyperparameter optimization through hypernetworks
J. Lorraine and D. Duvenaud · 2018
Later among the works it cites.
Truncated back-propagation for bilevel optimization
A. Shaban, C.-A. Cheng, N. Hatch, and B. Boots · 2018
Later among the works it cites.
Flipout: Efficient pseudo-independent weight perturbations on mini-batches
Y. Wen, P. Vicol, J. Ba, D. Tran, and R. Grosse · 2018
Later among the works it cites.
Randaugment: Practical data augmentation with no separate search
E. D. Cubuk, B. Zoph, J. Shlens, and Q. V. Le · 2019
Later among the works it cites.
Hyperparameter optimization
M. Feurer and F. Hutter · 2019
Later among the works it cites.
Convergence of learning dynamics in stackelberg games
T. Fiez, B. Chasnov, and L. J. Ratliff · 2019
Later among the works it cites.
What is local optimality in nonconvex-nonconcave minimax optimization?
C. Jin, P. Netrapalli, and M. I. Jordan · 2019
Later among the works it cites.
Optimizing millions of hyperparameters by implicit differentiation
J. Lorraine, P. Vicol, and D. Duvenaud · 2019
Later among the works it cites.
M. MacKay, P. Vicol, J. Lorraine, D. Duvenaud, and R. Grosse · 2019
Later among the works it cites.
On solving minimax optimization locally: A follow-the-ridge approach
Y. Wang, G. Zhang, and J. Ba · 2019
Later among the works it cites.