Fetching the paper…
Reading the bibliography…
We recapitulate the Bayesian formulation of neural network based classifiers and show that, while sampling from the posterior does indeed lead to better generalisation than is obtained by standard optimisation of the cost function, even better performance can in general be achieved by sampling finite temperature ($T$) distributions derived from the posterior.
“Equation of state calculations by fast computing machines”
Nicholas Metropolis et al · 1953
Earlier work this paper cites.
“Studies in molecular dynamics. I. General method”
Berni Alder and Thomas Wainwright · 1959
Earlier work this paper cites.
“A computer simulation method for the calculation of equilibrium constants for the formation of physical clusters of molecules: Application to small water clusters”
William Swope, Hans Andersen, Peter Berens and Kent Wilson · 1982
Earlier work this paper cites.
“Optimization by simulated annealing”
Scott Kirkpatrick, C Gelatt and Mario Vecchi · 1983
Earlier work this paper cites.
“New Monte Carlo method to compute the free energy of arbitrary solids. Application to the fcc and hcp phases of hard spheres”
Daan Frenkel and Anthony Ladd · 1984
Earlier work this paper cites.
“Replica Monte Carlo Simulation of Spin-Glasses”
Robert. Swendsen and Jian-Sheng Wang · 1986
Earlier work this paper cites.
“Hybrid monte carlo”
Simon Duane, Anthony Kennedy, Brian Pendleton and Duncan Roweth · 1987
Earlier work this paper cites.
“A generalized guided Monte Carlo algorithm”
Alan. Horowitz · 1991
Earlier work this paper cites.
“Bayesian Interpolation”
David.C. MacKay · 1991
Earlier work this paper cites.
“A practical Bayesian framework for backpropagation networks”
David MacKay · 1992
Earlier work this paper cites.
“The evidence framework applied to classification networks”
David MacKay · 1992
Earlier work this paper cites.
“Gradient-based learning applied to document recognition”
Yann LeCun, L“’eon Bottou, Yoshua Bengio and Patrick Haffner · 1998
Earlier work this paper cites.
“Replica-exchange molecular dynamics method for protein folding”
Yuji Sugita and Yuko Okamoto · 1999
Earlier work this paper cites.
“Multidimensional replica-exchange method for free-energy calculations”
Yuji Sugita, Akio Kitao and Yuko Okamoto · 2000
Cited alongside, same era.
“Understanding molecular simulation: from algorithms to applications”
Daan Frenkel and Berend Smit · 2001
Cited alongside, same era.
“Annealed importance sampling”
Radford Neal · 2001
Cited alongside, same era.
“On the acceptance probability of replica-exchange Monte Carlo trials”
David. Kofke · 2002
Cited alongside, same era.
“Structural Relaxation Made Simple”
Erik Bitzek et al · 2006
Cited alongside, same era.
“Feedback-optimized parallel tempering Monte Carlo”
Helmut Katzgraber, Simon Trebst, David Huse and Matthias Troyer · 2006
Cited alongside, same era.
“Bayesian learning via stochastic gradient Langevin dynamics”
Max Welling and Yee Teh · 2011
Later among the works it cites.
“Bayesian posterior sampling via stochastic gradient Fisher scoring”
Sungjin Ahn, Anoop Korattikara and Max Welling · 2012
Later among the works it cites.
“Bayesian learning for neural networks”
Radford Neal · 2012
Later among the works it cites.
“Replica exchange with Smart Monte Carlo and Hybrid Monte Carlo in manifolds”
R. Jenkins, E. Curotto and Massimo Mella · 2013
Later among the works it cites.
Guillaume Desjardins, Heng Luo, Aaron Courville and Yoshua Bengio · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Learning and evaluating Boltzmann machines”
Ruslan Salakhutdinov · 2008
Cited alongside, same era.
“On the quantitative analysis of deep belief networks”
Ruslan Salakhutdinov and Iain Murray · 2008
Cited alongside, same era.
“Parallel tempering is efficient for learning restricted Boltzmann machines”
KyungHyun Cho, Tapani Raiko and Alexander Ilin · 2010
Cited alongside, same era.
“Adaptive parallel tempering for stochastic maximum likelihood learning of RBMs”
Guillaume Desjardins, Aaron Courville and Yoshua Bengio · 2010
Cited alongside, same era.
“Metadynamics”
Alessandro Barducci, Massimiliano Bonomi and Michele Parrinello · 2011
Cited alongside, same era.
“Improved learning of Gaussian-Bernoulli restricted Boltzmann machines”
KyungHyun Cho, Alexander Ilin and Tapani Raiko · 2011
Cited alongside, same era.
Jascha Sohl-Dickstein, Mayur Mudigonda and Michael DeWeese · 2014
Later among the works it cites.
“Weight uncertainty in neural networks”
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu and Daan Wierstra · 2015
Later among the works it cites.
“Replica exchange Hybrid Monte Carlo simulations of the ammonia dodecamer and hexadecamer”
J.G. Venditto, S. Wolf, E. Curotto and Massimo Mella · 2015
Later among the works it cites.
“On the quantitative analysis of decoder-based generative models”
Yuhuai Wu, Yuri Burda, Ruslan Salakhutdinov and Roger Grosse · 2016
Later among the works it cites.
“Stochastic gradient descent as approximate bayesian inference”
Stephan Mandt, Matthew Hoffman and David Blei · 2017
Later among the works it cites.
“A bayesian perspective on generalization and stochastic gradient descent”
Samuel Smith and Quoc Le · 2017
Later among the works it cites.
“Langevin-gradient parallel tempering for Bayesian neural learning”
Rohitash Chandra, Konark Jain, Ratneel Deo and Sally Cripps · 2018
Later among the works it cites.
“nn_sample”
Robert.N. Baldock · 2019
Closest in time.