Fetching the paper…
Reading the bibliography…
Hyperparameter selection generally relies on running multiple full training trials, with selection based on validation set performance.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
Fast Exact Multiplication by the Hessian
Pearlmutter, B. A. (1994) · 1994
Earlier work this paper cites.
Adaptive regularization in neural network modeling
Larsen, J., Svarer, C., Andersen, L. N., and Hansen, L. K. (1998) · 1998
Earlier work this paper cites.
The MNIST database of handwritten digits
LeCun, Y., Cortes, C., and Burges, C. J. (1998) · 1998
Earlier work this paper cites.
Optimal use of regularization and cross-validation in neural network modeling
Chen, D. and Hagan, M. T. (1999) · 1999
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Schraudolph, N. N. (2002) · 2002
Earlier work this paper cites.
Efficient multiple hyperparameter learning for log-linear models
Foo, C.-s., Do, C. B., and Ng, A. (2008) · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. (2009) · 2009
Earlier work this paper cites.
Stacked denoising autoencoders: learning useful representations in a deep network with a local denoising criterion
Vincent, P., Larochelle, H., Lajoie, I., Bengio, Y., and antoine Manzagol, P. (2010) · 2010
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
Bergstra, J. S., Bardenet, R., Bengio, Y., and Kégl, B. (2011) · 2011
Cited alongside, same era.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. (2011) · 2011
Cited alongside, same era.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y. (2012) · 2012
Cited alongside, same era.
Deep learning made easier by linear transformations in perceptrons
Raiko, T., Valpola, H., and LeCun, Y. (2012) · 2012
Cited alongside, same era.
Practical Bayesian Optimization of Machine Learning Algorithms
Snoek, J., Larochelle, H., and Adams, R. P. (2012) · 2012
Cited alongside, same era.
Improving deep neural networks for LVCSR using rectified linear units and dropout
Dahl, G. E., Sainath, T. N., and Hinton, G. E. (2013) · 2013
Striving for Simplicity: The All Convolutional Net
Springenberg, J. T., Dosovitskiy, A., Brox, T., and Riedmiller, M. (2014) · 2014
Later among the works it cites.
Natural neural networks
Desjardins, G., Simonyan, K., Pascanu, R., and Kavukcuoglu, K. (2015) · 2015
Closest in time.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Closest in time.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J. (2015) · 2015
Closest in time.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, D., Duvenaud, D., and Adams, R. P. (2015) · 2015
Closest in time.
Semi-supervised learning with ladder network
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pushing stochastic gradient towards second-order methods–backpropagation learning with transformations in nonlinearities
Vatanen, T., Raiko, T., Valpola, H., and LeCun, Y. (2013) · 2013
Cited alongside, same era.
Fast dropout training
Wang, S. I. and Manning, C. D. (2013) · 2013
Cited alongside, same era.
Gradient-based optimization of hyperparameters
Bengio, Y. (2000)
Cited in the paper.
Rasmus, A., Valpola, H., Honkala, M., Berglund, M., and Raiko, T. (2015) · 2015
Closest in time.
Empirical evaluation of rectified activations in convolutional network
Xu, B., Wang, N., and Li, M. (2015) · 2015
Closest in time.
Theano: A Python framework for fast computation of mathematical expressions
Team, T. T. D. (2016) · 2016
Closest in time.