Sur la théorie du mouvement brownien
Langevin, P · 1908
Earlier work this paper cites.
Verification of forecasts expressed in terms of probability
Brier, G. W · 1950
Earlier work this paper cites.
On bias reduction in estimation
Schucany, W., Gray, H., and Owen, D · 1971
Earlier work this paper cites.
Statistical decision theory and Bayesian analysis
Berger, J. O · 1985
Earlier work this paper cites.
Replica Monte Carlo simulation of spin-glasses
Swendsen, R. H. and Wang, J.-S · 1986
Earlier work this paper cites.
Hybrid Monte Carlo
Duane, S., Kennedy, A. D., Pendleton, B. J., and Roweth, D · 1987
Earlier work this paper cites.
Inference from iterative simulation using multiple sequences
Gelman, A. and Rubin, D. B · 1992
Earlier work this paper cites.
An Introduction to Predictive Inference
Geisser, S · 1993
Earlier work this paper cites.
Keeping neural networks simple by minimizing the description length of the weights
Hinton, G. and Van Camp, D · 1993
Earlier work this paper cites.
Ensemble learning and evidence maximization
MacKay, D. J. et al · 1995
Earlier work this paper cites.
Bayesian learning for neural networks
Neal, R. M · 1995
Earlier work this paper cites.
Estimating functions of probability distributions from a finite set of samples
Wolpert, D. H. and Wolf, D. R · 1995
Earlier work this paper cites.
On asymptotic properties of predictive distributions
Komaki, F · 1996
Earlier work this paper cites.
Ensemble learning for multi-layer networks
Barber, D. and Bishop, C. M · 1998
Earlier work this paper cites.
Replica-exchange molecular dynamics method for protein folding
Sugita, Y. and Okamoto, Y · 1999
Earlier work this paper cites.
Entropy and inference, revisited
Nemenman, I., Shafee, F., and Bialek, W · 2002
Earlier work this paper cites.
On the Markov chain central limit theorem
Jones, G. L. et al · 2004
Earlier work this paper cites.
Parallel tempering: Theory, applications, and new perspectives
Earl, D. J. and Deem, M. W · 2005
Earlier work this paper cites.
Bootstrap prediction and Bayesian prediction under misspecified models
Fushiki, T. et al · 2005
Earlier work this paper cites.
On variance conditions for markov chain clts
Häggström, O. and Rosenthal, J · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Handbook of Markov Chain Monte Carlo
Brooks, S., Gelman, A., Jones, G., and Meng, X · 2011
Earlier work this paper cites.
MCMC using Hamiltonian dynamics
Neal, R. M. et al · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
Welling, M. and Teh, Y. W · 2011
Earlier work this paper cites.
Machine learning: a probabilistic perspective
Murphy, K. P · 2012
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
Bayesian data analysis
Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., and Rubin, D. B · 2013
Earlier work this paper cites.
Robust Bayesian inference under model misspecification, 2013
Jansen, L · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Earlier work this paper cites.
Stochastic gradient Hamiltonian Monte Carlo
Chen, T., Fox, E., and Guestrin, C · 2014
Earlier work this paper cites.
Bayesian sampling using stochastic gradient thermostats
Ding, N., Fang, Y., Babbush, R., Chen, C., Skeel, R. D., and Neven, H · 2014
Earlier work this paper cites.
The no-u-turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo
Hoffman, M. D. and Gelman, A · 2014
Earlier work this paper cites.