Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Jürgen Schmidhuber · 1987
Earlier work this paper cites.
Learning a synaptic learning rule
Yoshua Bengio, Samy Bengio, and Jocelyn Cloutier · 1990
Earlier work this paper cites.
Learning factorial codes by predictability minimization
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Samy Bengio, Yoshua Bengio, Jocelyn Cloutier, and Jan Gecsei · 1992
Earlier work this paper cites.
On learning how to learn learning strategies
Juergen Schmidhuber · 1995
Earlier work this paper cites.
An investigation of the gradient descent process in neural networks
Barak Pearlmutter · 1996
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
Bruno A Olshausen and David J Field · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Feature extraction through lococode
Sepp Hochreiter and Jürgen Schmidhuber · 1999
Earlier work this paper cites.
Evolution and design of distributed learning rules
Thomas Philip Runarsson and Magnus Thor Jonsson · 2000
Earlier work this paper cites.
Learning to learn using gradient descent
Sepp Hochreiter, A Steven Younger, and Peter R Conwell · 2001
Earlier work this paper cites.
A taxonomy of global optimization methods based on response surfaces
Donald R Jones · 2001
Earlier work this paper cites.
Evolving neural networks through augmenting topologies
Kenneth O Stanley and Risto Miikkulainen · 2002
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov · 2006
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, and Pierre-Antoine Manzagol · 2010
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning
Original
Quoc V Le, Marc’Aurelio Ranzato, Rajat Monga, Matthieu Devin, Kai Chen, Greg S Corrado, Jeff Dean, and Andrew Y Ng · 2011
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
James S Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Learning feature representations with k-means
Adam Coates and Andrew Y Ng · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.