Fetching the paper…
Reading the bibliography…
The key distinguishing property of a Bayesian approach is marginalization, rather than using a single setting of weights.
Bayesian inference in statistical analysis, addision-wesley
Box, G. E. and Tiao, G. C · 1973
Earlier work this paper cites.
On the ability of the optimal perceptron to generalise
Opper, M., Kinzel, W., Kleinz, J., and Nehl, R · 1990
Earlier work this paper cites.
Minimum complexity density estimation
Barron, A. R. and Cover, T. M · 1991
Earlier work this paper cites.
Bayesian methods for adaptive models
MacKay, D. J · 1992
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Hinton, G. E. and Van Camp, D · 1993
Earlier work this paper cites.
Bayes factors
Kass, R. E. and Raftery, A. E · 1995
Earlier work this paper cites.
Probable networks and plausible predictions?a review of practical Bayesian methods for supervised neural networks
MacKay, D. J · 1995
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Neal, R · 1996
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Introduction to Gaussian processes
MacKay, D. J · 1998
Earlier work this paper cites.
Adaptive and learning systems for signal processing communications, and control
Vapnik, V. N · 1998
Earlier work this paper cites.
Pac-bayesian model averaging
McAllester, D. A · 1999
Earlier work this paper cites.
Bayesian model averaging is not model combination
Minka, T. P · 2000
Earlier work this paper cites.
Occam’s razor
Rasmussen, C. E. and Ghahramani, Z · 2001
Earlier work this paper cites.
On Bayesian consistency
Walker, S. and Hjort, N. L · 2001
Earlier work this paper cites.
(not) bounding the true error
Langford, J. and Caruana, R · 2002
Earlier work this paper cites.
Information-theoretic upper and lower bounds for statistical estimation
Zhang, T · 2002
Earlier work this paper cites.
Variational algorithms for approximate Bayesian inference
Beal, M. J · 2003
Earlier work this paper cites.
Information theory, inference and learning algorithms
MacKay, D. J · 2003
Earlier work this paper cites.
Pattern Recognition and Machine Learning
Bishop, C. M · 2006
Earlier work this paper cites.
Gaussian processes for Machine Learning
Rasmussen, C. E. and Williams, C. K. I · 2006
Earlier work this paper cites.
Rademacher complexity bounds for non-iid processes
Mohri, M. and Rostamizadeh, A · 2009
Earlier work this paper cites.
Approximate Bayesian inference for latent gaussian models by using integrated nested laplace approximations
Rue, H., Martino, S., and Chopin, N · 2009
Earlier work this paper cites.
Bayesian methods for data-dependent priors
Darnieder, W. F · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Welling, M. and Teh, Y. W · 2011
Earlier work this paper cites.
The safe Bayesian
Grünwald, P · 2012
Earlier work this paper cites.
The CIFAR-10 dataset
Krizhevsky, A., Nair, V., and Hinton, G · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Cited alongside, same era.
Weight uncertainty in neural networks
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D · 2015
Cited alongside, same era.
A general framework for updating belief distributions
Bissiri, P. G., Holmes, C. C., and Walker, S. G · 2016
Cited alongside, same era.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2019
Later among the works it cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2019
Later among the works it cites.
Cobb, A. D., Baydin, A. G., Markham, A., and Roberts, S. J · 2019
Later among the works it cites.
Safe-Bayesian generalized linear regression
de Heide, R., Kirichenko, A., Mehta, N., and Grünwald, P · 2019
Later among the works it cites.
Large scale structure of neural network loss landscapes
Fort, S. and Jastrzebski, S · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Cited alongside, same era.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Cited alongside, same era.
Dziugaite, G. K. and Roy, D. M · 2017
Cited alongside, same era.
Inconsistency of Bayesian inference for misspecified linear models, and a proposal for repairing it
Grünwald, P., Van Ommen, T., et al · 2017
Cited alongside, same era.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q · 2017
Cited alongside, same era.
What uncertainties do we need in Bayesian deep learning for computer vision?
Kendall, A. and Gal, Y · 2017
Cited alongside, same era.
Deep ensembles: A loss landscape perspective
Fort, S., Hu, H., and Lakshminarayanan, B · 2019
Later among the works it cites.
A primer on pac-bayesian learning
Guedj, B · 2019
Later among the works it cites.
Evaluating scalable bayesian deep learning methods for robust computer vision
Gustafsson, F. K., Danelljan, M., and Schön, T. B · 2019
Later among the works it cites.
Benchmarking neural network robustness to common corruptions and perturbations
Hendrycks, D. and Dietterich, T · 2019
Later among the works it cites.
Understanding generalization through visualizations
Huang, W. R., Emam, Z., Goldblum, M., Fowl, L., Terry, J. K., Huang, F., and Goldstein, T · 2019
Later among the works it cites.
Subspace inference for Bayesian deep learning
Izmailov, P., Maddox, W. J., Kirichenko, P., Garipov, T., Vetrov, D., and Wilson, A. G · 2019
Later among the works it cites.
Fantastic generalization measures and where to find them
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D., and Bengio, S · 2019
Later among the works it cites.
The functional neural process
Louizos, C., Shi, X., Schutte, K., and Welling, M · 2019
Later among the works it cites.
A simple baseline for Bayesian uncertainty in deep learning
Maddox, W. J., Izmailov, P., Garipov, T., Vetrov, D. P., and Wilson, A. G · 2019
Later among the works it cites.
Learning under model misspecification: Applications to variational and ensemble methods, 2019
Masegosa, A. R · 2019
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 2019
Later among the works it cites.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Ovadia, Y., Fertig, E., Ren, J., Nado, Z., Sculley, D., Nowozin, S., Dillon, J. V., Lakshminarayanan, B., and Snoek, J · 2019
Later among the works it cites.
Functional variational Bayesian neural networks
Sun, S., Zhang, G., Shi, J., and Grosse, R · 2019
Later among the works it cites.
Calibrated bayes factors for model comparison
Xu, X., Lu, P., MacEachern, S., and Xu, R · 2019
Later among the works it cites.
Output-constrained Bayesian neural networks
Yang, W., Lorch, L., Graule, M. A., Srinivasan, S., Suresh, A., Yao, J., Pradier, M. F., and Doshi-Velez, F · 2019
Later among the works it cites.
Pitfalls of in-domain uncertainty estimation and ensembling in deep learning
Ashukha, A., Lyzhov, A., Molchanov, D., and Vetrov, D · 2020
Closest in time.
Optimal regularization can mitigate double descent
Nakkiran, P., Venkat, P., Kakade, S., and Ma, T · 2020
Closest in time.
Scipy 1.0: fundamental algorithms for scientific computing in python
Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., et al · 2020
Closest in time.
How good is the Bayes posterior in deep neural networks really?
Wenzel, F., Roth, K., Veeling, B. S., Światkowski, J., Tran, L., Mandt, S., Snoek, J., Salimans, T., Jenatton, R., and Nowozin, S · 2020
Closest in time.
The case for Bayesian deep learning
Wilson, A. G · 2020
Closest in time.
User-friendly introduction to pac-bayes bounds
Alquier, P · 2021
Closest in time.
What are Bayesian neural network posteriors really like?
Izmailov, P., Vikram, S., Hoffman, M. D., and Wilson, A. G · 2021
Closest in time.
On uncertainty, tempering, and data augmentation in bayesian classification
Kapoor, S., Maddox, W. J., Izmailov, P., and Wilson, A. G · 2022
Closest in time.