Fetching the paper…
Reading the bibliography…
Deep learning research aims at discovering learning algorithms that discover multiple levels of distributed representations, with higher levels representing more abstract concepts.
Markov Random Fields and Their Applications (Contemporary Mathematics ; V. 1)
Kindermann, R. (1980) · 1980
Earlier work this paper cites.
Optimization by simulated annealing
Kirkpatrick, S., Jr., C. D. G., , and Vecchi, M. P. (1983) · 1983
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D., Hinton, G., and Williams, R. (1986) · 1986
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen Netzen. Diploma thesis, Institut für Informatik, Lehrstuhl Prof. Brauer, Technische Universität München
Hochreiter, S. (1991) · 1991
Earlier work this paper cites.
A self-organizing neural network that discovers surfaces in random-dot stereograms
Becker, S. and Hinton, G. (1992) · 1992
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., and Frasconi, P. (1994) · 1994
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Neal, R. M. (1994) · 1994
Earlier work this paper cites.
The wake-sleep algorithm for unsupervised neural networks
Hinton, G. E., Dayan, P., Frey, B. J., and Neal, R. M. (1995) · 1995
Earlier work this paper cites.
Does the wake-sleep algorithm learn good density estimators?
Frey, B. J., Hinton, G. E., and Dayan, P. (1996) · 1996
Earlier work this paper cites.
Emergence of invariant-feature detectors in the adaptive-subspace self-organizing map
Kohonen, T. (1996) · 1996
Earlier work this paper cites.
Emergence of simple-cell receptive field properties by learning a sparse code for natural images
Olshausen, B. A. and Field, D. J. (1996) · 1996
Earlier work this paper cites.
Exploiting tractable substructures in intractable networks
Saul, L. K. and Jordan, M. I. (1996) · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Centering neural network gradient factors
Schraudolph, N. N. (1998) · 1998
Earlier work this paper cites.
Dependency networks for inference, collaborative filtering, and data visualization
Heckerman, D., Chickering, D. M., Meek, C., Rounthwaite, R., and Kadie, C. (2000) · 2000
Earlier work this paper cites.
Emergence of phase and shift invariant features by decomposition of natural images into independent feature subspaces
Hyvärinen, A. and Hoyer, P. (2000) · 2000
Earlier work this paper cites.
Separating style and content with bilinear models
Tenenbaum, J. B. and Freeman, W. T. (2000) · 2000
Earlier work this paper cites.
Quantum annealing of a disordered magnet
Brooke, J. J., Bitko, D., Rosenbaum, T. F., and Aeppli, G. (2001) · 2001
Earlier work this paper cites.
Extended ensemble monte carlo
Iba, Y. (2001) · 2001
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y., Ducharme, R., Vincent, P., and Jauvin, C. (2003) · 2003
Earlier work this paper cites.
Scaling large learning problems with hard parallel mixtures
Collobert, R., Bengio, Y., and Bengio., S. (2003) · 2003
Earlier work this paper cites.
Algorithms for manifold learning
Cayton, L. (2005) · 2005
Earlier work this paper cites.
Estimation of non-normalized statistical models using score matching
Hyvärinen, A. (2005) · 2005
Earlier work this paper cites.
Large margin methods for structured and interdependent output variables
Tsochantaridis, I., Joachims, T., Hofmann, T., and Altun, Y. (2005) · 2005
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Bengio, Y., Lamblin, P., Popovici, D., and Larochelle, H. (2007) · 2006
Earlier work this paper cites.
Pattern Recognition and Machine Learning
Bishop, C. M. (2006) · 2006
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Hinton, G. E. and Salakhutdinov, R. (2006) · 2006
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Hinton, G. E., Osindero, S., and Teh, Y. (2006) · 2006
Earlier work this paper cites.
A tutorial on energy-based learning
LeCun, Y., Chopra, S., Hadsell, R., Ranzato, M.-A., and Huang, F.-J. (2006) · 2006
Earlier work this paper cites.
Efficient learning of sparse representations with an energy-based model
Ranzato, M., Poultney, C., Chopra, S., and LeCun, Y. (2007) · 2006
Earlier work this paper cites.
Shift-invariant sparse coding for audio classification
Grosse, R., Raina, R., Kwong, H., and Ng, A. Y. (2007) · 2007
Earlier work this paper cites.
Echo state network
Jaeger, H. (2007) · 2007
Earlier work this paper cites.
Structured learning with approximate inference
Kulesza, A. and Pereira, F. (2008) · 2007
Earlier work this paper cites.
Self-taught learning: transfer learning from unlabeled data
Raina, R., Battle, A., Lee, H., Packer, B., and Ng, A. Y. (2007) · 2007
Earlier work this paper cites.
An introduction to quantum annelaing
Rose, G. and Macready, W. (2007) · 2007
Earlier work this paper cites.
Restricted Boltzmann machines for collaborative filtering
Salakhutdinov, R., Mnih, A., and Hinton, G. (2007) · 2007
Earlier work this paper cites.
Modeling human motion using binary latent variables
Taylor, G., Hinton, G. E., and Roweis, S. (2007) · 2007
Earlier work this paper cites.
Neural net language models
Bengio, Y. (2008) · 2008
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Collobert, R. and Weston, J. (2008) · 2008
Earlier work this paper cites.
Fast inference in sparse coding algorithms with applications to object recognition
Kavukcuoglu, K., Ranzato, M., and LeCun, Y. (2008) · 2008
Earlier work this paper cites.
Classification using discriminative restricted Boltzmann machines
Larochelle, H. and Bengio, Y. (2008) · 2008
Earlier work this paper cites.
Topmoumoute online natural gradient algorithm
Le Roux, N., Manzagol, P.-A., and Bengio, Y. (2008) · 2008
Earlier work this paper cites.
Sparse deep belief net model for visual area V2
Lee, H., Ekanadham, C., and Ng, A. (2008) · 2008
Earlier work this paper cites.
Sparse feature learning for deep belief networks
Ranzato, M., Boureau, Y.-L., and LeCun, Y. (2008) · 2008
Cited alongside, same era.
Extracting and composing robust features with denoising autoencoders
Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A. (2008) · 2008
Cited alongside, same era.
Deep learning via semi-supervised embedding
Weston, J., Ratle, F., and Collobert, R. (2008) · 2008
Cited alongside, same era.
Differentiable sparse coding
Bagnell, J. A. and Bradley, D. M. (2009) · 2009
Cited alongside, same era.
Learning deep architectures for AI
Bengio, Y. (2009) · 2009
Cited alongside, same era.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J. (2009) · 2009
Cited alongside, same era.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Recht, B., Re, C., Wright, S., and Niu, F. (2011) · 2011
Later among the works it cites.
Contractive auto-encoders: Explicit invariance during feature extraction
Rifai, S., Vincent, P., Muller, X., Glorot, X., and Bengio, Y. (2011a) · 2011
Later among the works it cites.
The manifold tangent classifier
Rifai, S., Dauphin, Y., Vincent, P., Bengio, Y., and Muller, X. (2011b) · 2011
Later among the works it cites.
Conversational speech transcription using context-dependent deep neural networks
Seide, F., Li, G., and Yu, D. (2011a) · 2011
Later among the works it cites.
Feature engineering in context-dependent deep neural networks for conversational speech transcription
Seide, F., Li, G., and Yu, D. (2011b) · 2011
Later among the works it cites.
Empirical risk minimization of graphical model parameters given approximate inference, decoding, and model structure
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bergstra, J. and Bengio, Y. (2009) · 2009
Cited alongside, same era.
Measuring invariances in deep networks
Goodfellow, I., Le, Q., Saxe, A., and Ng, A. (2009) · 2009
Cited alongside, same era.
What is the best multi-stage architecture for object recognition?
Jarrett, K., Kavukcuoglu, K., Ranzato, M., and LeCun, Y. (2009) · 2009
Cited alongside, same era.
Structured variable selection with sparsity-inducing norms
Jenatton, R., Audibert, J.-Y., and Bach, F. (2009) · 2009
Cited alongside, same era.
Online dictionary learning for sparse coding
Mairal, J., Bach, F., Ponce, J., and Sapiro, G. (2009) · 2009
Cited alongside, same era.
Large-scale deep unsupervised learning using graphics processors
Raina, R., Madhavan, A., and Ng, A. Y. (2009) · 2009
Cited alongside, same era.
Stoyanov, V., Ropson, A., and Eisner, J. (2011) · 2011
Later among the works it cites.
On autoencoders and score matching for energy based models
Swersky, K., Ranzato, M., Buchman, D., Marlin, B., and de Freitas, N. (2011) · 2011
Later among the works it cites.
A connection between score matching and denoising autoencoders
Vincent, P. (2011) · 2011
Later among the works it cites.
Bayesian learning via stochastic gradient Langevin dynamics
Welling, M. and Teh, Y.-W. (2011) · 2011
Later among the works it cites.
Learning image representations from the pixel level via hierarchical sparse coding
Yu, K., Lin, Y., and Lafferty, J. (2011) · 2011
Later among the works it cites.
Implicit density estimation by local moment matching to sample from auto-encoders
Bengio, Y., Alain, G., and Rifai, S. (2012) · 2012
Later among the works it cites.
Multi-column deep neural networks for image classification
Ciresan, D., Meier, U., and Schmidhuber, J. (2012) · 2012
Later among the works it cites.
Emergence of object-selective features in unsupervised feature learning
Coates, A., Karpathy, A., and Ng, A. (2012) · 2012
Later among the works it cites.
Deep networks for predicting ad click through rates
Corrado, G. (2012) · 2012
Later among the works it cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Le, Q., Mao, M., Ranzato, M., Senior, A., Tucker, P., Yang, K., and Ng, A. Y. (2012) · 2012
Later among the works it cites.
Disentangling factors of variation via generative entangling
Desjardins, G., Courville, A., and Bengio, Y. (2012) · 2012
Later among the works it cites.
Learning approximate inference policies for fast prediction
Eisner, J. (2012) · 2012
Later among the works it cites.
Large-scale feature learning with spike-and-slab sparse coding
Goodfellow, I., Courville, A., and Bengio, Y. (2012) · 2012
Later among the works it cites.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. (2012) · 2012
Later among the works it cites.
Building high-level features using large scale unsupervised learning
Le, Q., Ranzato, M., Monga, R., Devin, M., Corrado, G., Chen, K., Dean, J., and Ng, A. (2012) · 2012
Later among the works it cites.
Statistical Language Models based on Neural Networks
Mikolov, T. (2012) · 2012
Later among the works it cites.
Deep Boltzmann machines and the centering trick
Montavon, G. and Muller, K.-R. (2012) · 2012
Later among the works it cites.
Machine Learning: a Probabilistic Perspective
Murphy, K. P. (2012) · 2012
Later among the works it cites.
On the difficulty of training recurrent neural networks
Pascanu, R. and Bengio, Y. (2012) · 2012
Later among the works it cites.
Deep learning made easier by linear transformations in perceptrons
Raiko, T., Valpola, H., and LeCun, Y. (2012) · 2012
Later among the works it cites.
A generative process for sampling contractive auto-encoders
Rifai, S., Bengio, Y., Dauphin, Y., and Vincent, P. (2012b) · 2012
Later among the works it cites.
No More Pesky Learning Rates
Schaul, T., Zhang, S., and LeCun, Y. (2012) · 2012
Later among the works it cites.
Training Recurrent Neural Networks
Sutskever, I. (2012) · 2012
Later among the works it cites.
Communication/computation tradeoffs in consensus-based distributed optimization
Tsianos, K., Lawlor, S., and Rabbat, M. (2012) · 2012
Later among the works it cites.
What regularized auto-encoders learn from the data generating distribution
Alain, G. and Bengio, Y. (2013) · 2013
Closest in time.
Deep generative stochastic networks trainable by backprop
Bengio, Y. and Thibodeau-Laufer, E. (2013) · 2013
Closest in time.
Advances in optimizing recurrent networks
Bengio, Y., Boulanger-Lewandowski, N., and Pascanu, R. (2013a) · 2013
Closest in time.
Better mixing via deep representations
Bengio, Y., Mesnil, G., Dauphin, Y., and Rifai, S. (2013b) · 2013
Closest in time.
A semantic matching energy function for learning with multi-relational data
Bordes, A., Glorot, X., Weston, J., and Bengio, Y. (2013) · 2013
Closest in time.
Big neural networks waste capacity
Dauphin, Y. and Bengio, Y. (2013) · 2013
Closest in time.
Recent advances in deep learning for speech research at Microsoft
Deng, L., Li, J., Huang, J.-T., Yao, K., Yu, D., Seide, F., Seltzer, M., Zweig, G., He, X., Williams, J., Gong, Y., and Acero, A. (2013) · 2013
Closest in time.
Maxout networks
Goodfellow, I. J., Warde-Farley, D., Mirza, M., Courville, A., and Bengio, Y. (2013b) · 2013
Closest in time.
Knowledge matters: Importance of prior information for optimization
Gulcehre, C. and Bengio, Y. (2013) · 2013
Closest in time.
Exploring compositional high order pattern potentials for structured output learning
Li, Y., Tarlow, D., and Zemel, R. (2013) · 2013
Closest in time.
Texture modeling with convolutional spike-and-slab RBMs and deep extensions
Luo, H., Carrier, P. L., Courville, A., and Bengio, Y. (2013) · 2013
Closest in time.
Revisiting natural gradient for deep networks
Pascanu, R. and Bengio, Y. (2013) · 2013
Closest in time.
Learning and selecting features jointly with point-wise gated Boltzmann machines
Sohn, K., Zhou, G., and Lee, H. (2013) · 2013
Closest in time.
Stochastic pooling for regularization of deep convolutional neural networks
Zeiler, M. D. and Fergus, R. (2013) · 2013
Closest in time.