Fetching the paper…
Reading the bibliography…
Ensembles of deep neural networks demonstrate improved performance over single models.
Verification of forecasts expressed in terms of probability
Brier, G. W · 1950
Earlier work this paper cites.
Feasibility of multivariate density estimates
Scott, D. W · 1991
Earlier work this paper cites.
A practical Bayesian framework for backpropagation networks
MacKay, D. J · 1992
Earlier work this paper cites.
Bayesian learning for neural networks
Neal, R. M · 1996
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Schölkopf, B., Smola, A. J., and Bach, F · 2002
Earlier work this paper cites.
Optimal aggregation of classifiers in statistical learning
Tsybakov, A. B · 2004
Earlier work this paper cites.
Prior distributions for variance parameters in hierarchical models (comment on article by Browne and Draper)
Gelman, A · 2006
Earlier work this paper cites.
Gaussian Processes for Machine Learning
Rasmussen, C. E., Williams, C. K., and Bach, F · 2006
Earlier work this paper cites.
Gradient flows: in metric spaces and in the space of probability measures
Ambrosio, L., Gigli, N., and Savaré, G · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Optimal transport: old and new , volume 338
Villani, C · 2009
Earlier work this paper cites.
Practical variational inference for neural networks
Graves, A · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Welling, M. and Teh, Y. W · 2011
Earlier work this paper cites.
Weight uncertainty in neural network
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Kingma, D. P., Salimans, T., and Welling, M · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using Bayesian binning
Naeini, M. P., Cooper, G., and Hauskrecht, M · 2015
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N · 2015
Earlier work this paper cites.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Stein variational gradient descent: A general purpose Bayesian inference algorithm
Liu, Q. and Wang, D · 2016
Earlier work this paper cites.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2017
Cited alongside, same era.
Stein variational gradient descent as gradient flow
Liu, Q · 2017
Cited alongside, same era.
Stochastic gradient descent as approximate Bayesian inference
Mandt, S., Hoffman, M. D., and Blei, D. M · 2017
Cited alongside, same era.
Grad-CAM: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2017
Cited alongside, same era.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Allen-Zhu, Z. and Li, Y · 2020
Later among the works it cites.
Pitfalls of in-domain uncertainty estimation and ensembling in deep learning
Ashukha, A., Lyzhov, A., Molchanov, D., and Vetrov, D · 2020
Later among the works it cites.
Online knowledge distillation with diverse peers
Chen, D., Mei, J.-P., Wang, C., Feng, Y., and Chen, C · 2020
Later among the works it cites.
Projected Stein variational gradient descent
Chen, P. and Ghattas, O · 2020
Later among the works it cites.
Efficient and scalable Bayesian neural nets with rank-1 factors
Dusenberry, M., Jerfel, G., Wen, Y., Ma, Y., Snoek, J., Heller, K., Lakshminarayanan, B., and Tran, D · 2020
Later among the works it cites.
Stable behaviour of infinitely wide deep neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2017
Cited alongside, same era.
A unified particle-optimization framework for scalable Bayesian sampling
Chen, C., Zhang, R., Wang, W., Li, B., and Chen, L · 2018
Cited alongside, same era.
Stochastic gradient MCMC with repulsive forces
Gallego, V. and Insua, D. R · 2018
Cited alongside, same era.
Gradient estimators for implicit models
Li, Y. and Turner, R. E · 2018
Cited alongside, same era.
A scalable Laplace approximation for neural networks
Ritter, H., Botev, A., and Barber, D · 2018
Cited alongside, same era.
A spectral approach to gradient estimation for implicit distributions
Shi, J., Sun, S., and Zhu, J · 2018
Cited alongside, same era.
Favaro, S., Peluchetti, S., and Fortini, S · 2020
Later among the works it cites.
Shortcut learning in deep neural networks
Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A · 2020
Later among the works it cites.
Holes in Bayesian statistics
Gelman, A. and Yao, Y · 2020
Later among the works it cites.
Evaluating scalable Bayesian deep learning methods for robust computer vision
Gustafsson, F. K., Danelljan, M., and Schon, T. B · 2020
Later among the works it cites.
Subspace inference for Bayesian deep learning
Izmailov, P., Maddox, W. J., Kirichenko, P., Garipov, T., Vetrov, D., and Wilson, A. G · 2020
Later among the works it cites.
Distributionally robust neural networks
Sagawa, S., Koh, P. W., Hashimoto, T. B., and Liang, P · 2020
Later among the works it cites.
All you need is a good functional prior for Bayesian deep learning
Tran, B.-H., Rossi, S., Milios, D., and Filippone, M · 2020
Later among the works it cites.
BatchEnsemble: an alternative approach to efficient ensemble and lifelong learning
Wen, Y., Tran, D., and Ba, J · 2020
Later among the works it cites.
Bayesian deep learning and a probabilistic perspective of generalization
Wilson, A. G. and Izmailov, P · 2020
Later among the works it cites.
Repulsive deep ensembles are Bayesian
D’Angelo, F. and Fortuin, V · 2021
Later among the works it cites.
On Stein variational neural network ensembles
D’Angelo, F., Fortuin, V., and Wenzel, F · 2021
Later among the works it cites.
The role of permutation invariance in linear mode connectivity of neural networks
Entezari, R., Sedghi, H., Saukh, O., and Neyshabur, B · 2021
Later among the works it cites.
DICE: Diversity in deep ensembles via conditional redundancy adversarial estimation
Rame, A. and Cord, M · 2021
Later among the works it cites.
Projected Wasserstein gradient descent for high-dimensional Bayesian inference
Wang, Y., Chen, P., and Li, W · 2021
Later among the works it cites.
Last layer re-training is sufficient for robustness to spurious correlations
Kirichenko, P., Izmailov, P., and Wilson, A. G · 2022
Closest in time.
Agree to disagree: Diversity through disagreement for better transferability
Pagliardini, M., Jaggi, M., Fleuret, F., and Karimireddy, S. P · 2022
Closest in time.