Fetching the paper…
Reading the bibliography…
It has previously been hypothesized, and supported with some experimental evidence, that deeper representations, when well trained, tend to do a better job at disentangling the underlying factors of variation.
Almost optimal lower bounds for small depth circuits
Håstad, J. (1986) · 1986
Earlier work this paper cites.
Modèles connexionistes de l’apprentissage
LeCun, Y. (1987) · 1987
Earlier work this paper cites.
The development of the time-delay neural network architecture for speech recognition
Lang, K. J. and Hinton, G. E. (1988) · 1988
Earlier work this paper cites.
Generalization and network design strategies
LeCun, Y. (1989) · 1989
Earlier work this paper cites.
On the power of small-depth threshold circuits
Håstad, J. and Goldmann, M. (1991) · 1991
Earlier work this paper cites.
Autoencoders, minimum description length, and helmholtz free energy
Hinton, G. E. and Zemel, R. S. (1994) · 1993
Earlier work this paper cites.
Sampling from multimodal distributions using tempered transitions
Neal, R. M. (1994) · 1994
Earlier work this paper cites.
Learning many related tasks at the same time with backpropagation
Caruana, R. (1995) · 1995
Earlier work this paper cites.
A Bayesian/information theoretic model of learning via multiple task sampling
Baxter, J. (1997) · 1997
Earlier work this paper cites.
Gradient based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998) · 1998
Earlier work this paper cites.
Remapping somatosensory cortex after injury
Flor, H. (2003) · 2003
Earlier work this paper cites.
The curse of highly variable functions for local kernel machines
Bengio, Y., Delalleau, O., and Le Roux, N. (2006) · 2005
Cited alongside, same era.
Algorithms for manifold learning
Cayton, L. (2005) · 2005
Cited alongside, same era.
A fast learning algorithm for deep belief nets
Hinton, G. E., Osindero, S., and Teh, Y.-W. (2006) · 2006
Cited alongside, same era.
Scaling learning algorithms towards AI
Bengio, Y. and LeCun, Y. (2007) · 2007
Cited alongside, same era.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Collobert, R. and Weston, J. (2008) · 2008
Cited alongside, same era.
Learning deep architectures for AI
Bengio, Y. (2009) · 2009
Cited alongside, same era.
Sample complexity of testing the manifold hypothesis
Narayanan, H. and Mitter, S. (2010) · 2010
Later among the works it cites.
Learning deep Boltzmann machines using adaptive MCMC
Salakhutdinov, R. (2010a) · 2010
Later among the works it cites.
Learning in Markov random fields using tempered transitions
Salakhutdinov, R. (2010b) · 2010
Later among the works it cites.
Susskind, J., Anderson, A., and Hinton, G. E. (2010) · 2010
Later among the works it cites.
On the expressive power of deep architectures
Bengio, Y. and Delalleau, O. (2011) · 2011
Later among the works it cites.
Domain adaptation for large-scale sentiment classification: A deep learning approach
Glorot, X., Bordes, A., and Bengio, Y. (2011) · 2011
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Measuring invariances in deep networks
Goodfellow, I., Le, Q., Saxe, A., and Ng, A. (2009) · 2009
Cited alongside, same era.
Deep Boltzmann machines
Salakhutdinov, R. and Hinton, G. E. (2009) · 2009
Cited alongside, same era.
Parallel tempering is efficient for learning restricted boltzmann machines
Cho, K., Raiko, T., and Ilin, A. (2010) · 2010
Cited alongside, same era.
Tempered Markov chain monte carlo for training of restricted Boltzmann machine
Desjardins, G., Courville, A., Bengio, Y., Vincent, P., and Delalleau, O. (2010) · 2010
Cited alongside, same era.
Later among the works it cites.
Contracting auto-encoders: Explicit invariance during feature extraction
Rifai, S., Vincent, P., Muller, X., Glorot, X., and Bengio, Y. (2011a) · 2011
Later among the works it cites.
The manifold tangent classifier
Rifai, S., Dauphin, Y., Vincent, P., Bengio, Y., and Muller, X. (2011b) · 2011
Later among the works it cites.
A generative process for sampling contractive auto-encoders
Rifai, S., Bengio, Y., Dauphin, Y., and Vincent, P. (2012) · 2012
Closest in time.
Quickly generating representative samples from an rbm-derived process
Breuleux, O., Bengio, Y., and Vincent, P. (2011) · 2073
Closest in time.