Fetching the paper…
Reading the bibliography…
Learning algorithms for energy based Boltzmann architectures that rely on gradient descent are in general computationally prohibitive, typically due to the exponential number of terms involved in computing the partition function.
S. Geman, D. Geman, Stochastic Relaxation, Gibbs Distributions, and the Bayesian Restoration of Images, IEEE Transactions on Pattern Analysis and Machine Intelligence 6 (6) (1984) 721–741
1984
Earlier work this paper cites.
P. Smolensky, Information Processing in Dynamical Systems: Foundations of Harmony Theory, in: D. E. Rumelhart, J. L. McClelland (Eds.), Parallel Distributed Processing: Explorations in the Microstructure of Cognition (vol. 1), MIT Press, 1986, pp. 194–281
1986
Earlier work this paper cites.
D. E. Rumelhart, G. E. Hinton, R. J. Williams, Learning Internal Representations by Error Propagation, in: D. E. Rumelhart, J. L. McClelland (Eds.), Parallel Distributed Processing: Explorations in the Microstructure of Cognition (vol. 1), MIT Press, 1986, pp. 318–362
1986
Earlier work this paper cites.
D. J. C. MacKay, Failures of the one-step learning algorithm, unpublished Technical Report (2001)
2001
Earlier work this paper cites.
G. E. Hinton, Training Products of Experts by Minimizing Contrastive Divergence, Neural Computation 14 (2002) 1771–1800
2002
Earlier work this paper cites.
M. A. Carreira-Perpiñán, G. E. Hinton, On Contrastive Divergence Learning, in: International Workshop on Artificial Intelligence and Statistics, 2005, pp. 33–40
2005
Earlier work this paper cites.
A. Yuille, The Convergence of Contrastive Divergence, in: Advances in Neural Information Processing Systems (NIPS’04), Vol. 17, MIT Press, 2005, pp. 1593–1600
2005
Earlier work this paper cites.
A. C. Coolen, R. Kühn, P. Sollich, Theory of neural information processing systems, OUP Oxford, 2005
2005
Earlier work this paper cites.
G. E. Hinton, S. Osindero, Y. Teh, A Fast Learning Algorithm for Deep Belief Nets, Neural Computation 18 (7) (2006) 1527–1554
2006
Earlier work this paper cites.
G. E. Hinton, R. R. Salakhutdinov, Reducing the Dimensionality of Data with Neural Networks, Science 313 (5786) (2006) 504–507
2006
Earlier work this paper cites.
R. Salakhutdinov, A. Mnih, G. Hinton, Restricted Boltzmann Machines for Collaborative Filtering, in: Proceedings of the 24th international conference on Machine learning, ACM, 2007, pp. 791–798
2007
Earlier work this paper cites.
Y. Bengio, P. Lamblin, D. Popovici, H. Larochelle, Greedy Layer-wise Training of Deep Networks, in: Advances in Neural Information Processing (NIPS’06), Vol. 19, MIT Press, 2007, pp. 153–160
2007
Cited alongside, same era.
Y. Bengio, Y. LeCun, Chapter 14: Scaling Learning Algorithms towards AI, in: D. D. L. Bottou, O. Chapelle, J. Weston (Eds.), Large-Scale Kernel Machine, MIT Press, 2007, pp. 321–359
2007
Cited alongside, same era.
A. Hyvarinen, Connections between score matching, contrastive divergence, and pseudolikelihood for continuous-valued variables, IEEE Transactions on Neural Networks 18 (5) (2007) 1529–1531
2007
Cited alongside, same era.
N. Le Roux, Y. Bengio, Representational Power of Restricted Boltzmann Machines and Deep Belief Networks, Neural Computation 20 (6) (2008) 1631–1649
2008
Cited alongside, same era.
A. Fischer, C. Igel, Empirical Analysis of the Divergence of Gibbs Sampling Based Learning Algorithms for Restricted Boltzmann Machines, in: International Conference on Artificial Neural Networks (ICANN), Vol. 3, 2010, pp. 208–217
2010
Later among the works it cites.
G. Desjardins, A. Courville, Y. Bengio, P. Vincent, O. Delalleau, Parallel Tempering for Training of Restricted Boltzmann Machines, in: 13th International Conference on Artificial Intelligence and Statistics (AISTATS), 2010, pp. 145–152
2010
Later among the works it cites.
P. M. Long, R. A. Serveido, Restricted Boltzmann Machines are Hard to Approximately Evaluate or Simulate, in: International Conference on Machine Learning, 2010, pp. 703–710
2010
Later among the works it cites.
I. Sutskever, T. Tieleman, On the convergence properties of contrastive divergence, in: Y. W. Teh, M. Titterington (Eds.), Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, Vol. 9 of Proceedings of Machine Learning Research, PMLR, 2010, pp. 789–795
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Tieleman, Training Restricted Boltzmann Machines using Approximations to the Likelihood Gradient, in: 25th International Conference on Machine Learning, 2008, pp. 1064–1071
2008
Cited alongside, same era.
E. Farguell, F. Mazzanti, E. Gomez-Ramirez, Boltzmann Machines Reduction by High-order Decimation, IEEE Transactions on Neural Networks 19 (10) (2008) 1816–1821
2008
Cited alongside, same era.
Y. Bengio, Learning deep architectures for AI, Foundations and Trends in Machine Learning 2 (1) (2009) 1–127
2009
Cited alongside, same era.
H. Larochelle, Y. Bengio, J. Lourador, P. Lamblin, Exploring Strategies for Training Deep Neural Networks, Journal of Machine Learning Research 10 (2009) 1–40
2009
Cited alongside, same era.
Y. Bengio, O. Delalleau, Justifying and Generalizing Contrastive Divergence, Neural Computation 21 (6) (2009) 1601–1621
2009
Cited alongside, same era.
T. Tieleman, G. E. Hinton, Using Fast Weights to Improve Persistent Contrastive Divergence, in: 26th International Conference on Machine Learning, 2009, pp. 1033–1040
2009
Cited alongside, same era.
2010
Later among the works it cites.
A. Fischer, C. Igel, Bounding the Bias of Contrastive Divergence Learning, Neural Computation 23 (3) (2011) 664–673
2011
Later among the works it cites.
A.-R. Mohamed, G. E. Dahl, G. Hinton, Acoustic Modeling using Deep Belief Networks, IEEE Transactions on Audio, Speech, and Language Processing 20 (1) (2012) 14–22
2012
Later among the works it cites.
R. Karakida, M. Okada, S. ichi Amari, Dynamical analysis of contrastive divergence learning: Restricted boltzmann machines with gaussian visible units, Neural Networks 79 (2016) 78 – 87
2016
Later among the works it cites.
G. Carleo, M. Troyer, Solving the Quantum Many-Body Problem with Artificial Neural Networks, Science 355 (6325) (2017) 602–606
2017
Later among the works it cites.
O. Krause, A. Fischer, C. Igel, Population-Contrastive-Divergence: Does Consistency Help with RBM training?, Pattern Recognition Letters 102 (2018) 1–7
2018
Closest in time.
O. Breuleux, Y. Bengio, P. Vincent, Quickly Generating Representative Samples from an RBM-Derived Process, Neural Computation 23 (8) (2011) 2058–2073
2073
Closest in time.