Fetching the paper…
Reading the bibliography…
Stochastic neurons can be useful for a number of reasons in deep learning models, but in many cases they pose a challenging problem: how to estimate the gradient of a loss function with respect to the input of such stochastic neurons, i.e., can we "back-propagate" through these stochastic neurons? We examine this question, existing approaches, and present two novel families of solutions, applicable in different settings.
Boltzmann machines: Constraint satisfaction networks that learn
Hinton, G. E., Sejnowski, T. J., and Ackley, D. H. (1984) · 1984
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J. (1986) · 1986
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
Spall, J. C. (1992) · 1992
Earlier work this paper cites.
Hierarchical recurrent neural networks for long-term dependencies
El Hihi, S. and Bengio, Y. (1996) · 1996
Earlier work this paper cites.
Gradient learning in spiking neural networks by dynamic perturbations of conductances
Fiete, I. R. and Seung, H. S. (2006) · 2006
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Hinton, G. E., Osindero, S., and Teh, Y.-W. (2006) · 2006
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A. (2008) · 2008
Cited alongside, same era.
Learning deep architectures for AI
Bengio, Y. (2009) · 2009
Cited alongside, same era.
Semantic hashing
Salakhutdinov, R. and Hinton, G. (2009) · 2009
Cited alongside, same era.
Rectified linear units improve restricted Boltzmann machines
Nair, V. and Hinton, G. E. (2010) · 2010
Cited alongside, same era.
Deep sparse rectifier neural networks
Glorot, X., Bordes, A., and Bengio, Y. (2011) · 2011
Cited alongside, same era.
Neural networks for machine learning
Hinton, G. (2012) · 2012
Cited alongside, same era.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2012) · 2012
Later among the works it cites.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. (2012a) · 2012
Later among the works it cites.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. (2012b) · 2012
Later among the works it cites.
Deep learning of representations: Looking forward
Bengio, Y. (2013) · 2013
Closest in time.
Unsupervised feature learning and deep learning: A review and new perspectives
Bengio, Y., Courville, A., and Vincent, P. (2013) · 2013
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Maxout networks
Goodfellow, I. J., Warde-Farley, D., Mirza, M., Courville, A., and Bengio, Y. (2013) · 2013
Closest in time.