Fetching the paper…
Reading the bibliography…
When using deep, multi-layered architectures to build generative models of data, it is difficult to train all layers at once.
Maximum likelihood from incomplete data via the EM algorithm
A. P. Dempster, N. M. Laird, and D. B. Rubin · 1977
Earlier work this paper cites.
On the convergence properties of the EM algorithm
C. F. Jeff Wu · 1983
Earlier work this paper cites.
Information processing in dynamical systems: foundations of harmony theory
P. Smolensky · 1986
Earlier work this paper cites.
Auto-association by multilayer perceptrons and singular value decomposition
H. Bourlard and Y. Kamp · 1988
Earlier work this paper cites.
Learning stochastic feedforward networks
R. M. Neal · 1990
Earlier work this paper cites.
Bayesian back-propagation
Wray L. Buntine and Andreas S. Weigend · 1991
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Annealed importance sampling
Radford M. Neal · 1998
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
G.E. Hinton · 2002
Earlier work this paper cites.
Elements of information theory
Thomas M. Cover and Joy A. Thomas · 2006
Cited alongside, same era.
A fast learning algorithm for deep belief nets
G.E. Hinton, S. Osindero, and Yee-Whye Teh · 2006
Cited alongside, same era.
Reducing the dimensionality of data with neural networks
G.E. Hinton and R. Salakhutdinov · 2006
Cited alongside, same era.
Greedy layer-wise training of deep networks
Y. Bengio, P. Lamblin, V. Popovici, and H. Larochelle · 2007
Cited alongside, same era.
Scaling learning algorithms towards ai
Y. Bengio and Y. LeCun · 2007
Cited alongside, same era.
An empirical evaluation of deep architectures on problems with many factors of variation
H. Larochelle, D. Erhan, A. Courville, J. Bergstra, and Y. Bengio · 2007
Cited alongside, same era.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Later among the works it cites.
Justifying and generalizing contrastive divergence
Yoshua Bengio and Olivier Delalleau · 2009
Later among the works it cites.
Exploring strategies for training deep neural networks
H. Larochelle, Y. Bengio, J. Louradour, and P. Lamblin · 2009
Later among the works it cites.
Deep Boltzmann machines
Ruslan Salakhutdinov and Geoffrey Hinton · 2009
Later among the works it cites.
Algorithms for hyper-parameter optimization
James Bergstra, Rémy Bardenet, Yoshua Bengio, and Balázs Kégl · 2011
Later among the works it cites.
Unsupervised feature learning and deep learning: A review and new perspectives
Yoshua Bengio, Aaron C. Courville, and Pascal Vincent · 2012
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Representational power of restricted Boltzmann machines and deep belief networks
Nicolas Le Roux and Yoshua Bengio · 2008
Cited alongside, same era.
On the quantitative analysis of deep belief networks
Ruslan Salakhutdinov and Iain Murray · 2008
Cited alongside, same era.
Closest in time.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio · 2012
Closest in time.
A generative process for sampling contractive auto-encoders
Salah Rifai, Yoshua Bengio, Yann Dauphin, and Pascal Vincent · 2012
Closest in time.