Fetching the paper…
Reading the bibliography…
The success of machine learning algorithms generally depends on data representation, and we hypothesize that this is because different representations can entangle and hide more or less the different explanatory factors of variation behind the data.
Receptive fields of single neurons in the cat’s striate cortex
Hubel, D. H. and Wiesel, T. N. (1959) · 1959
Earlier work this paper cites.
Statistical analysis of non-lattice data
Besag, J. (1975) · 1975
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position
Fukushima, K. (1980) · 1980
Earlier work this paper cites.
Almost optimal lower bounds for small depth circuits
Håstad, J. (1986) · 1986
Earlier work this paper cites.
Learning distributed representations of concepts
Hinton, G. E. (1986) · 1986
Earlier work this paper cites.
Learning processes in an asymmetric threshold network
LeCun, Y. (1986) · 1986
Earlier work this paper cites.
Information processing in dynamical systems: Foundations of harmony theory
Smolensky, P. (1986) · 1986
Earlier work this paper cites.
Modèles connexionistes de l’apprentissage
LeCun, Y. (1987) · 1987
Earlier work this paper cites.
Auto-association by multilayer perceptrons and singular value decomposition
Bourlard, H. and Kamp, Y. (1988) · 1988
Earlier work this paper cites.
Generalization and network design strategies
LeCun, Y. (1989) · 1989
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W., and Jackel, L. D. (1989) · 1989
Earlier work this paper cites.
Document image defect models
Baird, H. (1990) · 1990
Earlier work this paper cites.
On the power of small-depth threshold circuits
Håstad, J. and Goldmann, M. (1991) · 1991
Earlier work this paper cites.
Blind separation of sources, part I: an adaptive algorithm based on neuromimetic architecture
Jutten, C. and Herault, J. (1991) · 1991
Earlier work this paper cites.
Tangent prop - A formalism for specifying selected invariances in an adaptive network
Simard, P., Victorri, B., LeCun, Y., and Denker, J. (1992) · 1991
Earlier work this paper cites.
A self-organizing neural network that discovers surfaces in random-dot stereograms
Becker, S. and Hinton, G. (1992) · 1992
Earlier work this paper cites.
Connectionist learning of belief networks
Neal, R. M. (1992) · 1992
Earlier work this paper cites.
A connectionist approach to speech recognition
Bengio, Y. (1993) · 1993
Earlier work this paper cites.
Autoencoders, minimum description length, and helmholtz free energy
Hinton, G. E. and Zemel, R. S. (1994) · 1993
Earlier work this paper cites.
Probabilistic inference using Markov chain Monte-Carlo methods
Neal, R. M. (1993) · 1993
Earlier work this paper cites.
Efficient pattern recognition using a new transformation distance
Simard, P. Y., LeCun, Y., and Denker, J. (1993) · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., and Frasconi, P. (1994) · 1994
Earlier work this paper cites.
Unsupervised learning of distributions on binary vectors using two layer networks
Freund, Y. and Haussler, D. (1994) · 1994
Earlier work this paper cites.
Emergence of simple-cell receptive field properties by learning a sparse code for natural images
Olshausen, B. A. and Field, D. J. (1996) · 1996
Earlier work this paper cites.
The independent components of natural scenes are edge filters
Bell, A. and Sejnowski, T. J. (1997) · 1997
Earlier work this paper cites.
EM algorithms for PCA and sensible PCA
Roweis, S. (1997) · 1997
Earlier work this paper cites.
Learning continuous attractors in recurrent networks
Seung, S. H. (1998) · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S. (1998) · 1998
Earlier work this paper cites.
Neural networks: tricks of the trade
Orr, G. and Muller, K.-R., editors (1998) · 1998
Earlier work this paper cites.
Nonlinear component analysis as a kernel eigenvalue problem
Schölkopf, B., Smola, A., and Müller, K.-R. (1998) · 1998
Earlier work this paper cites.
Products of experts
Hinton, G. E. (1999) · 1999
Earlier work this paper cites.
Object recognition from local scale invariant features
Lowe, D. (1999) · 1999
Earlier work this paper cites.
Hierarchical models of object recognition in cortex
Riesenhuber, M. and Poggio, T. (1999) · 1999
Earlier work this paper cites.
Probabilistic principal components analysis
Tipping, M. E. and Bishop, C. M. (1999) · 1999
Earlier work this paper cites.
On the convergence of Markovian stochastic algorithms with rapidly decreasing ergodicity rates
Younes, L. (1999) · 1999
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Hinton, G. E. (2000) · 2000
Earlier work this paper cites.
Emergence of phase and shift invariant features by decomposition of natural images into independent feature subspaces
Hyvärinen, A. and Hoyer, P. (2000) · 2000
Earlier work this paper cites.
Nonlinear dimensionality reduction by locally linear embedding
Roweis, S. and Saul, L. K. (2000) · 2000
Earlier work this paper cites.
A global geometric framework for nonlinear dimensionality reduction
Tenenbaum, J., de Silva, V., and Langford, J. C. (2000) · 2000
Earlier work this paper cites.
Charting a manifold
Brand, M. (2003) · 2002
Earlier work this paper cites.
Stochastic neighbor embedding
Hinton, G. E. and Roweis, S. (2003) · 2002
Earlier work this paper cites.
Temporal coherence, natural image sequences, and the visual cortex
Hurri, J. and Hyvärinen, A. (2003) · 2002
Earlier work this paper cites.
Manifold Parzen windows
Vincent, P. and Bengio, Y. (2003) · 2002
Earlier work this paper cites.
Learning sparse topographic representations with products of Student-t distributions
Welling, M., Hinton, G. E., and Osindero, S. (2003) · 2002
Earlier work this paper cites.
Slow feature analysis: Unsupervised learning of invariances
Wiskott, L. and Sejnowski, T. (2002) · 2002
Earlier work this paper cites.
Laplacian eigenmaps for dimensionality reduction and data representation
Belkin, M. and Niyogi, P. (2003) · 2003
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y., Ducharme, R., Vincent, P., and Jauvin, C. (2003) · 2003
Earlier work this paper cites.
Out-of-sample extensions for LLE, Isomap, MDS, Eigenmaps, and Spectral Clustering
Bengio, Y., Paiement, J.-F., Vincent, P., Delalleau, O., Le Roux, N., and Ouimet, M. (2004) · 2003
Earlier work this paper cites.
Hessian eigenmaps: new locally linear embedding techniques for high-dimensional data
Donoho, D. L. and Grimes, C. (2003) · 2003
Earlier work this paper cites.
Best practices for convolutional neural networks
Simard, D., Steinkraus, P. Y., and Platt, J. C. (2003) · 2003
Earlier work this paper cites.
Non-local manifold tangent learning
Bengio, Y. and Monperrus, M. (2005) · 2004
Earlier work this paper cites.
How are complex cell properties adapted to the statistics of natural stimuli?
Körding, K. P., Kayser, C., Einhäuser, W., and König, P. (2004) · 2004
Earlier work this paper cites.
Unsupervised learning of image manifolds by semidefinite programming
Weinberger, K. Q. and Saul, L. K. (2004) · 2004
Earlier work this paper cites.
The convergence of contrastive divergences
Yuille, A. L. (2005) · 2004
Earlier work this paper cites.
The curse of highly variable functions for local kernel machines
Bengio, Y., Delalleau, O., and Le Roux, N. (2006a) · 2005
Earlier work this paper cites.
Non-local manifold Parzen windows
Bengio, Y., Larochelle, H., and Vincent, P. (2006b) · 2005
Earlier work this paper cites.
Slow feature analysis yields a rich repertoire of complex cell properties
Berkes, P. and Wiskott, L. (2005) · 2005
Earlier work this paper cites.
On contrastive divergence learning
Carreira-Perpiñan, M. A. and Hinton, G. E. (2005) · 2005
Earlier work this paper cites.
Estimation of non-normalized statistical models using score matching
Hyvärinen, A. (2005) · 2005
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Bengio, Y., Lamblin, P., Popovici, D., and Larochelle, H. (2007) · 2006
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Hinton, G. E. and Salakhutdinov, R. (2006) · 2006
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Hinton, G. E., Osindero, S., and Teh, Y. (2006) · 2006
Earlier work this paper cites.
Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories
Lazebnik, S., Schmid, C., and Ponce, J. (2006) · 2006
Earlier work this paper cites.
Efficient learning of sparse representations with an energy-based model
Ranzato, M., Poultney, C., Chopra, S., and LeCun, Y. (2007) · 2006
Earlier work this paper cites.
Scaling learning algorithms towards AI
Bengio, Y. and LeCun, Y. (2007) · 2007
Earlier work this paper cites.
Shift-invariant sparse coding for audio classification
Grosse, R., Raina, R., Kwong, H., and Ng, A. Y. (2007) · 2007
Earlier work this paper cites.
Some extensions of score matching
Hyvärinen, A. (2007) · 2007
Earlier work this paper cites.
Echo state network
Jaeger, H. (2007) · 2007
Earlier work this paper cites.
Self-taught learning: transfer learning from unlabeled data
Raina, R., Battle, A., Lee, H., Packer, B., and Ng, A. Y. (2007) · 2007
Earlier work this paper cites.
Sparse feature learning for deep belief networks
Ranzato, M., Boureau, Y., and LeCun, Y. (2008) · 2007
Earlier work this paper cites.
Semantic hashing
Salakhutdinov, R. and Hinton, G. E. (2007) · 2007
Earlier work this paper cites.
Restricted Boltzmann machines for collaborative filtering
Salakhutdinov, R., Mnih, A., and Hinton, G. E. (2007) · 2007
Earlier work this paper cites.
Robust object recognition with cortex-like mechanisms
Serre, T., Wolf, L., Bileschi, S., and Riesenhuber, M. (2007) · 2007
Earlier work this paper cites.
Neural net language models
Bengio, Y. (2008) · 2008
Cited alongside, same era.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Collobert, R. and Weston, J. (2008) · 2008
Cited alongside, same era.
Empirical evaluation of convolutional RBMs for vision
Desjardins, G. and Bengio, Y. (2008) · 2008
Cited alongside, same era.
Optimal approximation of signal priors
Hyvärinen, A. (2008) · 2008
Cited alongside, same era.
Natural image denoising with convolutional networks
Jain, V. and Seung, S. H. (2008) · 2008
Cited alongside, same era.
Fast inference in sparse coding algorithms with applications to object recognition
Kavukcuoglu, K., Ranzato, M., and LeCun, Y. (2008) · 2008
Cited alongside, same era.
Deconvolutional networks
Zeiler, M., Krishnan, D., Taylor, G., and Fergus, R. (2010) · 2010
Later among the works it cites.
Structured sparsity through convex optimization
Bach, F., Jenatton, R., Mairal, J., and Obozinski, G. (2011) · 2011
Later among the works it cites.
Deep learning of representations for unsupervised and transfer learning
Bengio, Y. (2011) · 2011
Later among the works it cites.
On the expressive power of deep architectures
Bengio, Y. and Delalleau, O. (2011) · 2011
Later among the works it cites.
Algorithms for hyper-parameter optimization
Bergstra, J., Bardenet, R., Bengio, Y., and Kégl, B. (2011) · 2011
Later among the works it cites.
Ask the locals: multi-way local pooling for image recognition
Boureau, Y., Le Roux, N., Bach, F., Ponce, J., and LeCun, Y. (2011) · 2011
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Classification using discriminative restricted Boltzmann machines
Larochelle, H. and Bengio, Y. (2008) · 2008
Cited alongside, same era.
Sparse deep belief net model for visual area V2
Lee, H., Ekanadham, C., and Ng, A. (2008) · 2008
Cited alongside, same era.
Evaluating probabilities under high-dimensional latent variable models
Murray, I. and Salakhutdinov, R. (2009) · 2008
Cited alongside, same era.
Generative versus discriminative training of RBMs for classification of fMRI images
Schmah, T., Hinton, G. E., Zemel, R., Small, S. L., and Strother, S. (2009) · 2008
Cited alongside, same era.
The recurrent temporal restricted Boltzmann machine
Sutskever, I., Hinton, G., and Taylor, G. (2009) · 2008
Cited alongside, same era.
Training restricted Boltzmann machines using approximations to the likelihood gradient
Tieleman, T. (2008) · 2008
Cited alongside, same era.
Later among the works it cites.
Classification with scattering operators
Bruna, J. and Mallat, S. (2011) · 2011
Later among the works it cites.
Enhanced gradient and adaptive learning rate for training restricted Boltzmann machines
Cho, K., Raiko, T., and Ilin, A. (2011) · 2011
Later among the works it cites.
The importance of encoding versus training with sparse coding and vector quantization
Coates, A. and Ng, A. Y. (2011a) · 2011
Later among the works it cites.
Selecting receptive fields in deep networks
Coates, A. and Ng, A. Y. (2011b) · 2011
Later among the works it cites.
Natural language processing (almost) from scratch
Collobert, R., Weston, J., Bottou, L., Karlen, M., Kavukcuoglu, K., and Kuksa, P. (2011) · 2011
Later among the works it cites.
A spike and slab restricted Boltzmann machine
Courville, A., Bergstra, J., and Bengio, Y. (2011a) · 2011
Later among the works it cites.
Unsupervised models of images by spike-and-slab RBMs
Courville, A., Bergstra, J., and Bengio, Y. (2011b) · 2011
Later among the works it cites.
On tracking the partition function
Desjardins, G., Courville, A., and Bengio, Y. (2011) · 2011
Later among the works it cites.
Deep sparse rectifier neural networks
Glorot, X., Bordes, A., and Bengio, Y. (2011a) · 2011
Later among the works it cites.
Domain adaptation for large-scale sentiment classification: A deep learning approach
Glorot, X., Bordes, A., and Bengio, Y. (2011b) · 2011
Later among the works it cites.
Spike-and-slab sparse coding for unsupervised feature discovery
Goodfellow, I., Courville, A., and Bengio, Y. (2011) · 2011
Later among the works it cites.
Structured sparse coding via lateral inhibition
Gregor, K., Szlam, A., and LeCun, Y. (2011) · 2011
Later among the works it cites.
Should penalized least squares regression be interpreted as Maximum A Posteriori estimation?
Gribonval, R. (2011) · 2011
Later among the works it cites.
Temporal pooling and multiscale learning for automatic annotation and ranking of music audio
Hamel, P., Lemieux, S., Bengio, Y., and Eck, D. (2011) · 2011
Later among the works it cites.
Unsupervised learning of sparse features for scalable audio classification
Henaff, M., Jarrett, K., Kavukcuoglu, K., and LeCun, Y. (2011) · 2011
Later among the works it cites.
Transforming auto-encoders
Hinton, G., Krizhevsky, A., and Wang, S. (2011) · 2011
Later among the works it cites.
On optimization methods for deep learning
Le, Q., Ngiam, J., Coates, A., Lahiri, A., Prochnow, B., and Ng, A. (2011a) · 2011
Later among the works it cites.
ICA with reconstruction cost for efficient overcomplete feature learning
Le, Q. V., Karpenko, A., Ngiam, J., and Ng, A. Y. (2011b) · 2011
Later among the works it cites.
Learning hierarchical spatio-temporal features for action recognition with independent subspace analysis
Le, Q. V., Zou, W. Y., Yeung, S. Y., and Ng, A. Y. (2011c) · 2011
Later among the works it cites.
Asymptotic efficiency of deterministic estimators for discrete energy-based models: Ratio matching and pseudolikelihood
Marlin, B. and de Freitas, N. (2011) · 2011
Later among the works it cites.
Learning recurrent neural networks with Hessian-free optimization
Martens, J. and Sutskever, I. (2011) · 2011
Later among the works it cites.
Unsupervised and transfer learning challenge: a deep learning approach
Mesnil, G., Dauphin, Y., Glorot, X., Rifai, S., Bengio, Y., Goodfellow, I., Lavoie, E., Muller, X., Desjardins, G., Warde-Farley, D., Vincent, P., Courville, A., and Bergstra, J. (2011) · 2011
Later among the works it cites.
Empirical evaluation and combination of advanced language modeling techniques
Mikolov, T., Deoras, A., Kombrink, S., Burget, L., and Cernocky, J. (2011) · 2011
Later among the works it cites.
Learning deep energy models
Ngiam, J., Chen, Z., Koh, P., and Ng, A. (2011) · 2011
Later among the works it cites.
On deep generative models with applications to recognition
Ranzato, M., Susskind, J., Mnih, V., and Hinton, G. (2011) · 2011
Later among the works it cites.
Contractive auto-encoders: Explicit invariance during feature extraction
Rifai, S., Vincent, P., Muller, X., Glorot, X., and Bengio, Y. (2011a) · 2011
Later among the works it cites.
The manifold tangent classifier
Rifai, S., Dauphin, Y., Vincent, P., Bengio, Y., and Muller, X. (2011c) · 2011
Later among the works it cites.
Réseaux de neurones à relaxation entraînés par critère d’autoencodeur débruitant
Savard, F. (2011) · 2011
Later among the works it cites.
Conversational speech transcription using context-dependent deep neural networks
Seide, F., Li, G., and Yu, D. (2011a) · 2011
Later among the works it cites.
Feature engineering in context-dependent deep neural networks for conversational speech transcription
Seide, F., Li, G., and Yu, D. (2011b) · 2011
Later among the works it cites.
Dynamic pooling and unfolding recursive autoencoders for paraphrase detection
Socher, R., Huang, E. H., Pennington, J., Ng, A. Y., and Manning, C. D. (2011a) · 2011
Later among the works it cites.
Semi-supervised recursive autoencoders for predicting sentiment distributions
Socher, R., Pennington, J., Huang, E. H., Ng, A. Y., and Manning, C. D. (2011b) · 2011
Later among the works it cites.
Empirical risk minimization of graphical model parameters given approximate inference, decoding, and model structure
Stoyanov, V., Ropson, A., and Eisner, J. (2011) · 2011
Later among the works it cites.
On score matching for energy based models: Generalizing autoencoders and simplifying deep learning
Swersky, K., Ranzato, M., Buchman, D., Marlin, B., and de Freitas, N. (2011) · 2011
Later among the works it cites.
A connection between score matching and denoising autoencoders
Vincent, P. (2011) · 2011
Later among the works it cites.
Learning image representations from the pixel level via hierarchical sparse coding
Yu, K., Lin, Y., and Lafferty, J. (2011) · 2011
Later among the works it cites.
Unsupervised learning of visual invariance with temporal coherence
Zou, W. Y., Ng, A. Y., and Yu, K. (2011) · 2011
Later among the works it cites.
What regularized auto-encoders learn from the data generating distribution
Alain, G. and Bengio, Y. (2012) · 2012
Closest in time.
Implicit density estimation by local moment matching to sample from auto-encoders
Bengio, Y., Alain, G., and Rifai, S. (2012) · 2012
Closest in time.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y. (2012) · 2012
Closest in time.
Joint learning of words and meaning representations for open-text semantic parsing
Bordes, A., Glorot, X., Weston, J., and Bengio, Y. (2012) · 2012
Closest in time.
Modeling temporal dependencies in high-dimensional sequences: Application to polyphonic music generation and transcription
Boulanger-Lewandowski, N., Bengio, Y., and Vincent, P. (2012) · 2012
Closest in time.
Marginalized denoising autoencoders for domain adaptation
Chen, M., Xu, Z., Winberger, K. Q., and Sha, F. (2012) · 2012
Closest in time.
Multi-column deep neural networks for image classification
Ciresan, D., Meier, U., and Schmidhuber, J. (2012) · 2012
Closest in time.
Context-dependent pre-trained deep neural networks for large vocabulary speech recognition
Dahl, G. E., Yu, D., Deng, L., and Acero, A. (2012) · 2012
Closest in time.
On training deep Boltzmann machines
Desjardins, G., Courville, A., and Bengio, Y. (2012) · 2012
Closest in time.
How does the brain solve visual object recognition?
DiCarlo, J., Zoccolan, D., and Rust, N. (2012) · 2012
Closest in time.
Learning approximate inference policies for fast prediction
Eisner, J. (2012) · 2012
Closest in time.
Spike-and-slab sparse coding for unsupervised feature discovery
Goodfellow, I. J., Courville, A., and Bengio, Y. (2012) · 2012
Closest in time.
Deep neural networks for acoustic modeling in speech recognition
Hinton, G., Deng, L., Dahl, G. E., Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T., and Kingsbury, B. (2012) · 2012
Closest in time.
Multiple texture Boltzmann machines
Kivinen, J. J. and Williams, C. K. I. (2012) · 2012
Closest in time.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. (2012) · 2012
Closest in time.
Group invariant scattering
Mallat, S. (2012) · 2012
Closest in time.
Acoustic modeling using deep belief networks
Mohamed, A., Dahl, G., and Hinton, G. (2012) · 2012
Closest in time.
When does a mixture of products contain a product of mixtures?
Montufar, G. F. and Morton, J. (2012) · 2012
Closest in time.
Deep learning made easier by linear transformations in perceptrons
Raiko, T., Valpola, H., and LeCun, Y. (2012) · 2012
Closest in time.
A generative process for sampling contractive auto-encoders
Rifai, S., Bengio, Y., Dauphin, Y., and Vincent, P. (2012) · 2012
Closest in time.
Large, pruned or continuous space language models on a gpu for statistical machine translation
Schwenk, H., Rousseau, A., and Attik, M. (2012) · 2012
Closest in time.
Practical bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P. (2012) · 2012
Closest in time.
Multimodal learning with deep boltzmann machines
Srivastava, N. and Salakhutdinov, R. (2012) · 2012
Closest in time.
Training Recurrent Neural Networks
Sutskever, I. (2012) · 2012
Closest in time.
Practical recommendations for gradient-based training of deep architectures
Bengio, Y. (2013) · 2013
Closest in time.
Better mixing via deep representations
Bengio, Y., Mesnil, G., Dauphin, Y., and Rifai, S. (2013) · 2013
Closest in time.
Structured output layer neural network language models for speech recognition
Le, H.-S., Oparin, I., Allauzen, A., Gauvin, J.-L., and Yvon, F. (2013) · 2013
Closest in time.
Pascanu, R. and Bengio, Y. (2013) · 2013
Closest in time.
Quickly generating representative samples from an RBM-derived process
Breuleux, O., Bengio, Y., and Vincent, P. (2011) · 2073
Closest in time.