Fetching the paper…
Reading the bibliography…
Learning algorithms related to artificial neural networks and in particular for Deep Learning may seem to involve many bells and whistles, called hyper-parameters.
A stochastic approximation method
Robbins, H. and Monro, S. (1951) · 1951
Earlier work this paper cites.
Relaxation and its role in vision
Hinton, G. E. (1978) · 1978
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovski, A. and Yudin, D. (1983) · 1983
Earlier work this paper cites.
Almost optimal lower bounds for small depth circuits
Håstad, J. (1986) · 1986
Earlier work this paper cites.
Learning distributed representations of concepts
Hinton, G. E. (1986) · 1986
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J. (1986) · 1986
Earlier work this paper cites.
Modèles connexionistes de l’apprentissage
LeCun, Y. (1987) · 1987
Earlier work this paper cites.
The development of the time-delay neural network architecture for speech recognition
Lang, K. J. and Hinton, G. E. (1988) · 1988
Earlier work this paper cites.
Connectionist learning procedures
Hinton, G. E. (1989) · 1989
Earlier work this paper cites.
Generalization and network design strategies
LeCun, Y. (1989) · 1989
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W., and Jackel, L. D. (1989) · 1989
Earlier work this paper cites.
Indexing by latent semantic analysis
Deerwester, S., Dumais, S. T., Furnas, G. W., Landauer, T. K., and Harshman, R. (1990) · 1990
Earlier work this paper cites.
Recursive distributed representations
Pollack, J. B. (1990) · 1990
Earlier work this paper cites.
On the power of small-depth threshold circuits
Håstad, J. and Goldmann, M. (1991) · 1991
Earlier work this paper cites.
Neural networks and the bias/variance dilemma
Geman, S., Bienenstock, E., and Doursat, R. (1992) · 1992
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. and Juditsky, A. (1992) · 1992
Earlier work this paper cites.
Multitask connectionist learning
Caruana, R. (1993) · 1993
Earlier work this paper cites.
Learning and development in neural networks: The importance of starting small
Elman, J. L. (1993) · 1993
Earlier work this paper cites.
Bagging predictors
Breiman, L. (1994) · 1994
Earlier work this paper cites.
Fast exact multiplication by the Hessian
Pearlmutter, B. (1994) · 1994
Earlier work this paper cites.
Learning internal representations
Baxter, J. (1995) · 1995
Earlier work this paper cites.
A Bayesian/information theoretic model of learning via multiple task sampling
Baxter, J. (1997) · 1997
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: a strategy employed by V1?
Olshausen, B. A. and Field, D. J. (1997) · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S. (1998) · 1998
Earlier work this paper cites.
A general framework for adaptive processing of data structures
Frasconi, P., Gori, M., and Sperduti, A. (1998) · 1998
Earlier work this paper cites.
Centering neural network gradient factors
Schraudolph, N. N. (1998) · 1998
Earlier work this paper cites.
A global geometric framework for nonlinear dimensionality reduction
Tenenbaum, J., de Silva, V., and Langford, J. C. (2000) · 2000
Earlier work this paper cites.
Applying slow feature analysis to image sequences yields a rich repertoire of complex cell properties
Berkes, P. and Wiskott, L. (2002) · 2002
Earlier work this paper cites.
Slow feature analysis: Unsupervised learning of invariances
Wiskott, L. and Sejnowski, T. J. (2002) · 2002
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y., Ducharme, R., Vincent, P., and Jauvin, C. (2003) · 2003
Earlier work this paper cites.
Large-scale on-line learning
Bottou, L. and LeCun, Y. (2004) · 2003
Earlier work this paper cites.
Links between perceptrons, MLPs and SVMs
Collobert, R. and Bengio, S. (2004a) · 2004
Earlier work this paper cites.
Convex neural networks
Bengio, Y., Le Roux, N., Vincent, P., Delalleau, O., and Marcotte, P. (2006a) · 2005
Earlier work this paper cites.
The curse of highly variable functions for local kernel machines
Bengio, Y., Delalleau, O., and Le Roux, N. (2006b) · 2005
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Bengio, Y., Lamblin, P., Popovici, D., and Larochelle, H. (2007) · 2006
Earlier work this paper cites.
Introduction to Statistical Relational Learning
Getoor, L. and Taskar, B. (2006) · 2006
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Hinton, G. E., Osindero, S., and Teh, Y.-W. (2006) · 2006
Earlier work this paper cites.
Markov logic networks
Richardson, M. and Domingos, P. (2006) · 2006
Earlier work this paper cites.
Scaling learning algorithms towards AI
Bengio, Y. and LeCun, Y. (2007) · 2007
Cited alongside, same era.
Efficient learning of sparse representations with an energy-based model
Ranzato, M., Poultney, C., Chopra, S., and LeCun, Y. (2007) · 2007
Cited alongside, same era.
Sparse feature learning for deep belief networks
Ranzato, M., Boureau, Y., and LeCun, Y. (2008b) · 2007
Cited alongside, same era.
Neural net language models
Bengio, Y. (2008) · 2008
Cited alongside, same era.
The tradeoffs of large scale learning
Bottou, L. and Bousquet, O. (2008) · 2008
Cited alongside, same era.
Classification using discriminative restricted Boltzmann machines
Larochelle, H. and Bengio, Y. (2008) · 2008
Cited alongside, same era.
Deep learning of representations for unsupervised and transfer learning
Bengio, Y. (2011) · 2011
Later among the works it cites.
On the expressive power of deep architectures
Bengio, Y. and Delalleau, O. (2011) · 2011
Later among the works it cites.
Algorithms for hyper-parameter optimization
Bergstra, J., Bardenet, R., Bengio, Y., and Kégl, B. (2011) · 2011
Later among the works it cites.
Learning structured embeddings of knowledge bases
Bordes, A., Weston, J., Collobert, R., and Bengio, Y. (2011) · 2011
Later among the works it cites.
From machine learning to machine reasoning
Bottou, L. (2011) · 2011
Later among the works it cites.
Enhanced gradient and adaptive learning rate for training restricted boltzmann machines
Cho, K., Raiko, T., and Ilin, A. (2011) · 2011
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Le Roux, N., Manzagol, P.-A., and Bengio, Y. (2008) · 2008
Cited alongside, same era.
Sparse deep belief net model for visual area V2
Lee, H., Ekanadham, C., and Ng, A. (2008) · 2008
Cited alongside, same era.
Visualizing data using t-sne
van der Maaten, L. and Hinton, G. E. (2008) · 2008
Cited alongside, same era.
Extracting and composing robust features with denoising autoencoders
Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A. (2008) · 2008
Cited alongside, same era.
Deep learning via semi-supervised embedding
Weston, J., Ratle, F., and Collobert, R. (2008) · 2008
Cited alongside, same era.
Differentiable sparse coding
Bagnell, J. A. and Bradley, D. M. (2009) · 2009
Cited alongside, same era.
The importance of encoding versus training with sparse coding and vector quantization
Coates, A. and Ng, A. Y. (2011) · 2011
Later among the works it cites.
Unsupervised models of images by spike-and-slab RBMs
Courville, A., Bergstra, J., and Bengio, Y. (2011) · 2011
Later among the works it cites.
Sampled reconstruction for large-scale learning of embeddings
Dauphin, Y., Glorot, X., and Bengio, Y. (2011) · 2011
Later among the works it cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Later among the works it cites.
Deep sparse rectifier neural networks
Glorot, X., Bordes, A., and Bengio, Y. (2011a) · 2011
Later among the works it cites.
Domain adaptation for large-scale sentiment classification: A deep learning approach
Glorot, X., Bordes, A., and Bengio, Y. (2011b) · 2011
Later among the works it cites.
Spike-and-slab sparse coding for unsupervised feature discovery
Goodfellow, I., Courville, A., and Bengio, Y. (2011) · 2011
Later among the works it cites.
Sequential model-based optimization for general algorithm configuration
Hutter, F., Hoos, H., and Leyton-Brown, K. (2011) · 2011
Later among the works it cites.
On optimization methods for deep learning
Le, Q., Ngiam, J., Coates, A., Lahiri, A., Prochnow, B., and Ng, A. (2011) · 2011
Later among the works it cites.
Improving first and second-order methods by modeling uncertainty
Le Roux, N., Bengio, Y., and Fitzgibbon, A. (2011) · 2011
Later among the works it cites.
Unsupervised and transfer learning challenge: a deep learning approach
Mesnil, G., Dauphin, Y., Glorot, X., Rifai, S., Bengio, Y., Goodfellow, I., Lavoie, E., Muller, X., Desjardins, G., Warde-Farley, D., Vincent, P., Courville, A., and Bergstra, J. (2011) · 2011
Later among the works it cites.
Contracting auto-encoders: Explicit invariance during feature extraction
Rifai, S., Vincent, P., Muller, X., Glorot, X., and Bengio, Y. (2011a) · 2011
Later among the works it cites.
The manifold tangent classifier
Rifai, S., Dauphin, Y., Vincent, P., Bengio, Y., and Muller, X. (2011b) · 2011
Later among the works it cites.
On random weights and unsupervised feature learning
Saxe, A. M., Koh, P. W., Chen, Z., Bhand, M., Suresh, B., and Ng, A. (2011) · 2011
Later among the works it cites.
Parsing natural scenes and natural language with recursive neural networks
Socher, R., Manning, C., and Ng, A. Y. (2011) · 2011
Later among the works it cites.
Parameter screening and optimisation for ILP using designed experiments
Srinivasan, A. and Ramakrishnan, G. (2011) · 2011
Later among the works it cites.
A connection between score matching and denoising autoencoders
Vincent, P. (2011) · 2011
Later among the works it cites.
Wsabie: Scaling up to large vocabulary image annotation
Weston, J., Bengio, S., and Usunier, N. (2011) · 2011
Later among the works it cites.
Unsupervised learning of visual invariance with temporal coherence
Zou, W. Y., Ng, A. Y., and Yu, K. (2011) · 2011
Later among the works it cites.
Implicit density estimation by local moment matching to sample from auto-encoders
Bengio, Y., Alain, G., and Rifai, S. (2012) · 2012
Closest in time.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y. (2012) · 2012
Closest in time.
Joint learning of words and meaning representations for open-text semantic parsing
Bordes, A., Glorot, X., Weston, J., and Bengio, Y. (2012) · 2012
Closest in time.
Le Roux, N., Schmidt, M., and Bach, F. (2012) · 2012
Closest in time.
Deep boltzmann machines as feed-forward hierarchies
Montavon, G., Braun, M. L., and Muller, K.-R. (2012) · 2012
Closest in time.
Deep learning made easier by linear transformations in perceptrons
Raiko, T., Valpola, H., and LeCun, Y. (2012) · 2012
Closest in time.
A generative process for sampling contractive auto-encoders
Rifai, S., Bengio, Y., Dauphin, Y., and Vincent, P. (2012) · 2012
Closest in time.
No More Pesky Learning Rates
Schaul, T., Zhang, S., and LeCun, Y. (2012) · 2012
Closest in time.
Large-scale learning with stochastic gradient descent
Bottou, L. (2013) · 2013
Closest in time.
A practical guide to training restricted boltzmann machines
Hinton, G. E. (2013) · 2013
Closest in time.
to appear
LeCun, Y. (2013) · 2013
Closest in time.
Quickly generating representative samples from an rbm-derived process
Breuleux, O., Bengio, Y., and Vincent, P. (2011) · 2073
Closest in time.