Fetching the paper…
Reading the bibliography…
The key distinguishing property of a Bayesian approach is marginalization instead of optimization, not the prior, or Bayes rule.
Functional variational bayesian neural networks
Sun, S., Zhang, G., Shi, J., and Grosse, R. (2019) · 1903
Earlier work this paper cites.
Output-constrained Bayesian neural networks
Yang, W., Lorch, L., Graule, M. A., Srinivasan, S., Suresh, A., Yao, J., Pradier, M. F., and Doshi-Velez, F. (2019) · 1905
Earlier work this paper cites.
Evaluating scalable Bayesian deep learning methods for robust computer vision
Gustafsson, F. K., Danelljan, M., and Schön, T. B. (2019) · 1906
Earlier work this paper cites.
Understanding generalization through visualizations
Huang, W. R., Emam, Z., Goldblum, M., Fowl, L., Terry, J. K., Huang, F., and Goldstein, T. (2019) · 1906
Earlier work this paper cites.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Ovadia, Y., Fertig, E., Ren, J., Nado, Z., Sculley, D., Nowozin, S., Dillon, J. V., Lakshminarayanan, B., and Snoek, J. (2019) · 1906
Earlier work this paper cites.
Probable networks and plausible predictions?a review of practical bayesian methods for supervised neural networks
MacKay, D. J. (1995) · 1995
Earlier work this paper cites.
Fractional Bayes factors for model comparison
O’Hagan, A. (1995) · 1995
Earlier work this paper cites.
The intrinsic Bayes factor for model selection and prediction
Berger, J. O. and Pericchi, L. R. (1996) · 1996
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Neal, R. (1996) · 1996
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Bayesian model averaging is not model combination
Minka, T. P. (2000) · 2000
Earlier work this paper cites.
Information theory, inference and learning algorithms
MacKay, D. J. (2003) · 2003
Earlier work this paper cites.
Model uncertainty
Clyde, M. and George, E. I. (2004) · 2004
Earlier work this paper cites.
The case for objective Bayesian analysis
Berger, J. et al. (2006) · 2006
Cited alongside, same era.
Bayesian modelling in machine learning: A tutorial review
Seeger, M. (2006) · 2006
Cited alongside, same era.
Gaussian processes for machine learning
Williams, C. K. and Rasmussen, C. E. (2006) · 2006
Cited alongside, same era.
Bayesian data analysis
Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., and Rubin, D. B. (2013) · 2013
Cited alongside, same era.
Covariance kernels for fast automatic pattern discovery and extrapolation with Gaussian processes
Wilson, A. G. (2014) · 2014
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z. (2016) · 2016
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of DNNs
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G. (2018) · 2018
Later among the works it cites.
Reliable uncertainty estimates in deep neural networks using noise contrastive priors
Hafner, D., Tran, D., Irpan, A., Lillicrap, T., and Davidson, J. (2018) · 2018
Later among the works it cites.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G. (2018) · 2018
Later among the works it cites.
Fast and scalable bayesian deep learning by weight-perturbation in adam
Khan, M. E., Nielsen, D., Tangkaratt, V., Lin, W., Gal, Y., and Srivastava, A. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P. (2016) · 2016
Cited alongside, same era.
Deep kernel learning
Wilson, A. G., Hu, Z., Salakhutdinov, R., and Xing, E. P. (2016) · 2016
Cited alongside, same era.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y. (2017) · 2017
Cited alongside, same era.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017) · 2017
Cited alongside, same era.
What uncertainties do we need in Bayesian deep learning for computer vision?
Kendall, A. and Gal, Y. (2017) · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C. (2017) · 2017
Cited alongside, same era.
Pradier, M. F., Pan, W., Yao, J., Ghosh, S., and Doshi-Velez, F. (2018) · 2018
Later among the works it cites.
A scalable Laplace approximation for neural networks
Ritter, H., Botev, A., and Barber, D. (2018) · 2018
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. (2018) · 2018
Later among the works it cites.
Subspace inference for Bayesian deep learning
Izmailov, P., Maddox, W. J., Kirichenko, P., Garipov, T., Vetrov, D., and Wilson, A. G. (2019) · 2019
Later among the works it cites.
The functional neural process
Louizos, C., Shi, X., Schutte, K., and Welling, M. (2019) · 2019
Later among the works it cites.
A simple baseline for Bayesian uncertainty in deep learning
Maddox, W., Garipov, T., Izmailov, P., Vetrov, D., and Wilson, A. G. (2019) · 2019
Later among the works it cites.
Classifier-agnostic saliency map extraction
Zołna, K., Geras, K. J., and Cho, K. (2019) · 2019
Later among the works it cites.
Cyclical stochastic gradient MCMC for Bayesian deep learning
Zhang, R., Li, C., Zhang, J., Chen, C., and Wilson, A. G. (2020) · 2020
Closest in time.