Fetching the paper…
Reading the bibliography…
We investigate the cold posterior effect through the lens of PAC-Bayes generalization bounds.
Merging of opinions with increasing information
D. Blackwell and L. Dubins · 1962
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Elements of information theory
T. M. Cover · 1999
Earlier work this paper cites.
Some PAC-Bayesian theorems
D. A. McAllester · 1999
Earlier work this paper cites.
Convergence rates of posterior distributions
S. Ghosal, J. K. Ghosh, and A. W. Van Der Vaart · 2000
Earlier work this paper cites.
PAC-Bayesian supervised classification: the thermodynamics of statistical learning , volume 56 of Monograph Series
O. Catoni · 2007
Earlier work this paper cites.
Suboptimal behavior of Bayes and MDL in classification under misspecification
P. Grünwald and J. Langford · 2007
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
L. Deng · 2012
Earlier work this paper cites.
The safe Bayesian
P. Grünwald · 2012
Earlier work this paper cites.
Weight uncertainty in neural network
C. Blundell, J. Cornebise, K. Kavukcuoglu, and D. Wierstra · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
M. P. Naeini, G. Cooper, and M. Hauskrecht · 2015
Earlier work this paper cites.
On the properties of variational approximations of Gibbs posteriors
P. Alquier, J. Ridgway, and N. Chopin · 2016
Earlier work this paper cites.
Pac-bayesian bounds based on the rényi divergence
L. Bégin, P. Germain, F. Laviolette, and J.-F. Roy · 2016
Earlier work this paper cites.
PAC-Bayesian theory meets Bayesian inference
P. Germain, F. Bach, A. Lacoste, and S. Lacoste-Julien · 2016
Earlier work this paper cites.
Emnist: Extending mnist to handwritten letters
G. Cohen, S. Afshar, J. Tapson, and A. Van Schaik · 2017
Earlier work this paper cites.
UCI machine learning repository, 2017
D. Dua and C. Graff · 2017
Earlier work this paper cites.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
G. K. Dziugaite and D. M. Roy · 2017
Cited alongside, same era.
Inconsistency of Bayesian inference for misspecified linear models, and a proposal for repairing it
P. Grünwald and T. Van Ommen · 2017
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Cited alongside, same era.
Fast and scalable bayesian deep learning by weight-perturbation in adam
M. Khan, D. Nielsen, V. Tangkaratt, W. Lin, Y. Gal, and A. Srivastava · 2018
Cited alongside, same era.
Slang: Fast structured covariance approximations for bayesian deep learning with natural gradient
A. Mishkin, F. Kunstner, D. Nielsen, M. Schmidt, and M. E. Khan · 2018
Cited alongside, same era.
How good is the Bayes posterior in deep neural networks really?
F. Wenzel, K. Roth, B. S. Veeling, J. Swiatkowski, L. Tran, S. Mandt, J. Snoek, T. Salimans, R. Jenatton, and S. Nowozin · 2020
Later among the works it cites.
Predicting training time without training
L. Zancato, A. Achille, A. Ravichandran, R. Bhotika, and S. Soatto · 2020
Later among the works it cites.
Why cold posteriors? on the suboptimal generalization of optimal Bayes estimates
C. Zeno, I. Golan, A. Pakman, and D. Soudry · 2020
Later among the works it cites.
Correlated input-dependent label noise in large-scale image classification
M. Collier, B. Mustafa, E. Kokiopoulou, R. Jenatton, and J. Berent · 2021
Later among the works it cites.
Laplace redux-effortless bayesian deep learning
E. Daxberger, A. Kristiadi, A. Immer, R. Eschenhagen, M. Bauer, and P. Hennig · 2021
Later among the works it cites.
A linearized framework and a new benchmark for model selection for fine-tuning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A scalable Laplace approximation for neural networks
H. Ritter, A. Botev, and D. Barber · 2018
Cited alongside, same era.
Pitfalls of in-domain uncertainty estimation and ensembling in deep learning
A. Ashukha, A. Lyzhov, D. Molchanov, and D. Vetrov · 2019
Cited alongside, same era.
Backpack: Packing more into backprop
F. Dangel, F. Kunstner, and P. Hennig · 2019
Cited alongside, same era.
Approximate inference turns deep networks into Gaussian processes
M. E. E. Khan, A. Immer, E. Abedi, and M. Korzepa · 2019
Cited alongside, same era.
Limitations of the empirical fisher approximation for natural gradient descent
F. Kunstner, P. Hennig, and L. Balles · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Cited alongside, same era.
Cold posteriors and aleatoric uncertainty
B. Adlam, J. Snoek, and S. L. Smith · 2020
Cited alongside, same era.
A. Deshpande, A. Achille, A. Ravichandran, H. Li, L. Zancato, C. Fowlkes, R. Bhotika, S. Soatto, and P. Perona · 2021
Later among the works it cites.
On the role of data in pac-bayes
G. K. Dziugaite, K. Hsu, W. Gharbieh, G. Arpino, and D. Roy · 2021
Later among the works it cites.
Bayesian neural network priors revisited
V. Fortuin, A. Garriga-Alonso, F. Wenzel, G. Rätsch, R. Turner, M. van der Wilk, and L. Aitchison · 2021
Later among the works it cites.
Improving predictions of bayesian neural nets via local linearization
A. Immer, M. Korzepa, and M. Bauer · 2021
Later among the works it cites.
What are Bayesian neural network posteriors really like?
P. Izmailov, S. Vikram, M. D. Hoffman, and A. G. G. Wilson · 2021
Later among the works it cites.
Bayesian confidence calibration for epistemic uncertainty modelling
F. Küppers, J. Kronenberger, J. Schneider, and A. Haselhoff · 2021
Later among the works it cites.
Fast adaptation with linearized neural networks
W. Maddox, S. Tang, P. Moreno, A. G. Wilson, and A. Damianou · 2021
Later among the works it cites.
Data augmentation in bayesian neural networks and the cold posterior effect
S. Nabarro, S. Ganev, A. Garriga-Alonso, V. Fortuin, M. van der Wilk, and L. Aitchison · 2021
Later among the works it cites.
Disentangling the roles of curation, data-augmentation and the prior in the cold posterior effect
L. Noci, K. Roth, G. Bachmann, S. Nowozin, and T. Hofmann · 2021
Later among the works it cites.
Speedy performance estimation for neural architecture search
R. Ru, C. Lyle, L. Schut, M. Fil, M. van der Wilk, and Y. Gal · 2021
Later among the works it cites.
Bayesian model selection, the marginal likelihood, and generalization
S. Lotfi, P. Izmailov, G. Benton, M. Goldblum, and A. G. Wilson · 2022
Closest in time.