A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Modeling by shortest data description
Jorma Rissanen · 1978
Earlier work this paper cites.
Stochastic complexity and modeling
Jorma Rissanen · 1986
Earlier work this paper cites.
Arithmetic coding for data compression
Ian H Witten, Radford M Neal, and John G Cleary · 1987
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S Denker, Sara A Solla, Richard E Howard, and Lawrence D Jackel · 1989
Earlier work this paper cites.
Classification by minimum-message-length inference
Chris S Wallace · 1990
Earlier work this paper cites.
Simplifying neural networks by soft weight-sharing
Steven J Nowlan and Geoffrey E Hinton · 1992
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Geoffrey E Hinton and Drew Van Camp · 1993
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Variational learning and bits-back coding: an information-theoretic view to bayesian learning
Antti Honkela and Harri Valpola · 2004
Earlier work this paper cites.
Pattern recognition
Christopher M Bishop · 2006
Earlier work this paper cites.
Practical variational inference for neural networks
Alex Graves · 2011
Earlier work this paper cites.
Multiresolution mixture modeling using merging of mixture components
Prem Raj Adhikari and Jaakko Hollmén · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.