Fetching the paper…
Reading the bibliography…
The "cold posterior effect" (CPE) in Bayesian deep learning describes the uncomforting observation that the predictive performance of Bayesian neural networks can be significantly improved if the Bayes posterior is artificially sharpened using a temperature parameter T<1.
Partitioned integrators for thermodynamic parameterization of neural networks
Leimkuhler, B., Matthews, C., and Vlaar, T. (2019) · 1908
Earlier work this paper cites.
A practical bayesian framework for backpropagation networks
MacKay, D. J. (1992) · 1992
Earlier work this paper cites.
Keeping neural networks simple by minimizing the description length of the weights
Hinton, G. and Van Camp, D. (1993) · 1993
Earlier work this paper cites.
The helmholtz machine
Dayan, P., Hinton, G. E., Neal, R. M., and Zemel, R. S. (1995) · 1995
Earlier work this paper cites.
Ensemble learning and evidence maximization
MacKay, D. J. et al. (1995) · 1995
Earlier work this paper cites.
Bayesian learning for neural networks
Neal, R. M. (1995) · 1995
Earlier work this paper cites.
Ensemble learning for multi-layer networks
Barber, D. and Bishop, C. M. (1998) · 1998
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020) · 2001
Earlier work this paper cites.
Bayesian deep learning and a probabilistic perspective of generalization
Wilson, A. G. and Izmailov, P. (2020) · 2002
Earlier work this paper cites.
Cold posteriors and aleatoric uncertainty
Adlam, B., Snoek, J., and Smith, S. L. (2020) · 2008
Earlier work this paper cites.
A statistical theory of cold posteriors in deep neural networks
Aitchison, L. (2020) · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G. (2009) · 2009
Earlier work this paper cites.
Dataset shift in machine learning
Quionero-Candela, J., Sugiyama, M., Schwaighofer, A., and Lawrence, N. D. (2009) · 2009
Earlier work this paper cites.
Safe learning: bridging the gap between bayes, mdl and statistical learning theory via empirical convexity
Grünwald, P. (2011) · 2011
Cited alongside, same era.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. (2011) · 2011
Cited alongside, same era.
Bayesian learning via stochastic gradient Langevin dynamics
Welling, M. and Teh, Y. W. (2011) · 2011
Cited alongside, same era.
The safe bayesian
Grünwald, P. (2012) · 2012
Cited alongside, same era.
Robust Bayesian inference under model misspecification
Jansen, L. (2013) · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G. (2013) · 2013
Cited alongside, same era.
Inconsistency of bayesian inference for misspecified linear models, and a proposal for repairing it
Grünwald, P., Van Ommen, T., et al. (2017) · 2017
Later among the works it cites.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017) · 2017
Later among the works it cites.
Cyclical stochastic gradient MCMC for Bayesian deep learning
Zhang, R., Li, C., Zhang, J., Chen, C., and Wilson, A. G. (2020) · 2017
Later among the works it cites.
Atanov, A., Ashukha, A., Struminsky, K., Vetrov, D., and Welling, M. (2018) · 2018
Later among the works it cites.
Learning invariances using the marginal likelihood
Wilk, M. v. d., Bauer, M., John, S., and Hensman, J. (2018) · 2018
Later among the works it cites.
Bayesian fractional posteriors
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stochastic gradient Hamiltonian Monte Carlo
Chen, T., Fox, E., and Guestrin, C. (2014) · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Cited alongside, same era.
Weight uncertainty in neural network
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D. (2015) · 2015
Cited alongside, same era.
A complete recipe for stochastic gradient MCMC
Ma, Y.-A., Chen, T., and Fox, E. (2015) · 2015
Cited alongside, same era.
Bridging the gap between stochastic gradient mcmc and stochastic optimization
Chen, C., Carlson, D., Gan, Z., Li, C., and Carin, L. (2016) · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Bhattacharya, A., Pati, D., Yang, Y., et al. (2019) · 2019
Later among the works it cites.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Ovadia, Y., Fertig, E., Ren, J., Nado, Z., Sculley, D., Nowozin, S., Dillon, J. V., Lakshminarayanan, B., and Snoek, J. (2019) · 2019
Later among the works it cites.
Human uncertainty makes classification more robust
Peterson, J. C., Battleday, R. M., Griffiths, T. L., and Russakovsky, O. (2019) · 2019
Later among the works it cites.
The case for Bayesian deep learning
Wilson, A. G. (2019) · 2019
Later among the works it cites.
How good is the bayes posterior in deep neural networks really?
Wenzel, F., Roth, K., Veeling, B., Swiatkowski, J., Tran, L., Mandt, S., Snoek, J., Salimans, T., Jenatton, R., and Nowozin, S. (2020) · 2020
Later among the works it cites.
Why cold posteriors? on the suboptimal generalization of optimal bayes estimates
Zeno, C., Golan, I., Pakman, A., and Soudry, D. (2020) · 2020
Later among the works it cites.
Bayesian neural network priors revisited
Fortuin, V., Garriga-Alonso, A., Wenzel, F., Rätsch, G., Turner, R., van der Wilk, M., and Aitchison, L. (2021) · 2021
Closest in time.
What are bayesian neural network posteriors really like?
Izmailov, P., Vikram, S., Hoffman, M. D., and Wilson, A. G. (2021) · 2021
Closest in time.