Fetching the paper…
Reading the bibliography…
For discrete data, the likelihood $P(x)$ can be rewritten exactly and parametrized into $P(X = x) = P(X = x | H = f(x)) P(H = f(x))$ if $P(X | H)$ has enough capacity to put no probability mass on any $x'$ for which $f(x')\neq f(x)$, where $f(\cdot)$ is a deterministic discrete function.
Connectionist learning procedures
Hinton, G. E. (1989) · 1989
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Bengio, Y., Lamblin, P., Popovici, D., and Larochelle, H. (2007) · 2006
Earlier work this paper cites.
Reducing the Dimensionality of Data with Neural Networks
Hinton, G. E. and Salakhutdinov, R. (2006) · 2006
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Hinton, G. E., Osindero, S., and Teh, Y.-W. (2006) · 2006
Earlier work this paper cites.
Learning deep architectures for AI
Bengio, Y. (2009) · 2009
Earlier work this paper cites.
Theano: a CPU and GPU math expression compiler
Bergstra, J., Breuleux, O., Bastien, F., Lamblin, P., Pascanu, R., Desjardins, G., Turian, J., Warde-Farley, D., and Bengio, Y. (2010) · 2010
Earlier work this paper cites.
Deep learning via Hessian-free optimization
Martens, J. (2010) · 2010
Cited alongside, same era.
Theano: new features and speed improvements
Bastien, F., Lamblin, P., Pascanu, R., Bergstra, J., Goodfellow, I. J., Bergeron, A., Bouchard, N., and Bengio, Y. (2012) · 2012
Cited alongside, same era.
Neural networks for machine learning
Hinton, G. (2012) · 2012
Cited alongside, same era.
Estimating or propagating gradients through stochastic neurons
Bengio, Y. (2013) · 2013
Cited alongside, same era.
How auto-encoders could provide credit assignment in deep networks via target propagation
Bengio, Y. (2014) · 2014
Cited alongside, same era.
Deep autoregressive networks
Gregor, K., Danihelka, I., Mnih, A., Blundell, C., and Wierstra, D. (2014) · 2014
Closest in time.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M. (2014) · 2014
Closest in time.
Neural variational inference and learning in belief networks
Mnih, A. and Gregor, K. (2014) · 2014
Closest in time.
A deep and tractable density estimator
Murray, B. U. I. and Larochelle, H. (2014) · 2014
Closest in time.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D. (2014) · 2014
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bornschein, J. and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Better mixing via deep representations
Bengio, Y., Mesnil, G., Dauphin, Y., and Rifai, S. (2013a)
Cited in the paper.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A. (2013b)
Cited in the paper.