A methodology for LISP program construction from examples
Phillip D. Summers · 1977
Earlier work this paper cites.
The inference of regular LISP programs from examples
Alan W. Biermann · 1978
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Juergen Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Object recognition from local scale-invariant features
David G. Lowe · 1999
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin · 2003
Earlier work this paper cites.
Histograms of oriented gradients for human detection
Navneet Dalal and Bill Triggs · 2005
Earlier work this paper cites.
Modeling systems with internal state using evolino
Daan Wierstra, Faustino J Gomez, and Jürgen Schmidhuber · 2005
Earlier work this paper cites.
Three new graphical models for statistical language modelling
Andriy Mnih and Geoffrey Hinton · 2007
Earlier work this paper cites.
Neuroevolution: from architectures to learning
Dario Floreano, Peter Dürr, and Claudio Mattiussi · 2008
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
Kevin Jarrett, Koray Kavukcuoglu, Yann Lecun, et al · 2009
Earlier work this paper cites.
A hypercube-based encoding for evolving large-scale neural networks
Kenneth O. Stanley, David B. D’Ambrosio, and Jason Gauci · 2009
Earlier work this paper cites.
Learning programs: A hierarchical Bayesian approach
Percy Liang, Michael I. Jordan, and Dan Klein · 2010
Earlier work this paper cites.
Rectified linear units improve restricted Boltzmann machines
Vinod Nair and Geoffrey E. Hinton · 2010
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
James Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio · 2012
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V. Le, et al · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E. Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N. Sainath, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Context dependent recurrent neural network language model
Tomas Mikolov and Geoffrey Zweig · 2012
Earlier work this paper cites.