Sur les opérations dans les ensembles abstraits et leur application aux équations intégrales
Stefan Banach · 1922
Earlier work this paper cites.
On Bayesian methods for seeking the extremum
Jonas Močkus · 1975
Earlier work this paper cites.
Hedonic housing prices and the demand for clean air
David Harrison Jr and Daniel L Rubinfeld · 1978
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: The meta-meta-… hook
Jürgen Schmidhuber · 1987
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Design and regularization of neural networks: The optimal use of a validation set
Jan Larsen, Lars Kai Hansen, Claus Svarer, and M Ohlsson · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, Patrick Haffner, et al · 1998
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Yoshua Bengio · 2000
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Generic methods for optimization-based modeling
Justin Domke · 2012
Earlier work this paper cites.
Practical Bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Neural networks for machine learning. Lecture 6a. Overview of mini-batch gradient descent
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky · 2012
Earlier work this paper cites.
Regularization of neural networks using Dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann Le Cun, and Rob Fergus · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Earlier work this paper cites.
Bilevel optimization with nonsmooth lower level problems
Peter Ochs, René Ranftl, Thomas Brox, and Thomas Pock · 2015
Earlier work this paper cites.
Optimizing neural networks with Kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
U-Net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Durk P Kingma, Tim Salimans, and Max Welling · 2015
Earlier work this paper cites.
Scalable gradient-based tuning of continuous regularization hyperparameters
Jelena Luketina, Mathias Berglund, Klaus Greff, and Tapani Raiko · 2016
Earlier work this paper cites.