Fetching the paper…
Reading the bibliography…
We propose SWA-Gaussian (SWAG), a simple, scalable, and general purpose approach for uncertainty representation and calibration in deep learning.
Cyclical stochastic gradient mcmc for bayesian deep learning
Zhang, R., Li, C., Zhang, J., Chen, C., and Wilson, A. G. (2019) · 1902
Earlier work this paper cites.
Subspace inference for bayesian deep learning
Izmailov, P., Maddox, W. J., Kirichenko, P., Garipov, T., Vetrov, D., and Wilson, A. G. (2019) · 1907
Earlier work this paper cites.
Efficient Estimators from a Slowly Convergent Robbins-Munro Process
Ruppert, D. (1988) · 1988
Earlier work this paper cites.
Acceleration of Stochastic Approximation by Averaging
Polyak, B. T. and Juditsky, A. B. (1992) · 1992
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Neal, R. M. (1996) · 1996
Earlier work this paper cites.
Asymptotic Statistics
Vaart, A. W. v. d. (1998) · 1998
Earlier work this paper cites.
Information theory, inference, and learning algorithms
MacKay, D. J. C. (2003) · 2003
Earlier work this paper cites.
Predicting good probabilities with supervised learning
Niculescu-Mizil, A. and Caruana, R. (2005) · 2005
Earlier work this paper cites.
Stochastic simulation: algorithms and analysis
Asmussen, S. and Glynn, P. W. (2007) · 2007
Earlier work this paper cites.
On-Line Estimation with the Multivariate Gaussian Distribution
Dasgupta, S. and Hsu, D. (2007) · 2007
Earlier work this paper cites.
An Analysis of Single-Layer Networks in Unsupervised Feature Learning
Coates, A., Ng, A., and Lee, H. (2011) · 2011
Earlier work this paper cites.
Practical variational inference for neural networks
Graves, A. (2011) · 2011
Earlier work this paper cites.
Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions
Halko, N., Martinsson, P.-G., and Tropp, J. A. (2011) · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Welling, M. and Teh, Y. W. (2011) · 2011
Earlier work this paper cites.
Statistical decision theory and Bayesian analysis
Berger, J. O. (2013) · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M. (2013) · 2013
Earlier work this paper cites.
Risk of bayesian inference in misspecified models, and the sandwich covariance matrix
Müller, U. K. (2013) · 2013
Earlier work this paper cites.
Stochastic Gradient Hamiltonian Monte Carlo
Chen, T., Fox, E. B., and Guestrin, C. (2014) · 2014
Earlier work this paper cites.
Weight Uncertainty in Neural Networks
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D. (2015) · 2015
Earlier work this paper cites.
Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks
Hernández-Lobato, J. M. and Adams, R. (2015) · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Cited alongside, same era.
Obtaining well calibrated probabilities using bayesian binning
Naeini, M. P., Cooper, G. F., and Hauskrecht, M. (2015) · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L. (2015) · 2015
Cited alongside, same era.
Deep gaussian processes for regression using approximate expectation propagation
Bui, T., Hernández-Lobato, D., Hernandez-Lobato, J., Li, Y., and Turner, R. (2016) · 2016
Cited alongside, same era.
Statistical Inference for Model Parameters in Stochastic Gradient Descent
Chen, X., Lee, J. D., Tong, X. T., and Zhang, Y. (2016) · 2016
Cited alongside, same era.
Constant step size stochastic gradient descent for probabilistic modeling
Babichev, D. and Bach, F. (2018) · 2018
Later among the works it cites.
The Description Length of Deep Learning models
Blier, L. and Ollivier, Y. (2018) · 2018
Later among the works it cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Chaudhari, P. and Soatto, S. (2018) · 2018
Later among the works it cites.
Essentially No Barriers in Neural Network Energy Landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F. A. (2018) · 2018
Later among the works it cites.
Gpytorch: Blackbox matrix-matrix gaussian process inference with gpu acceleration
Gardner, J., Pleiss, G., Weinberger, K. Q., Bindel, D., and Wilson, A. G. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dropout as a Bayesian Approximation
Gal, Y. and Ghahramani, Z. (2016) · 2016
Cited alongside, same era.
Deep Residual Learning for Image Recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Wide Residual Networks
Zagoruyko, S. and Komodakis, N. (2016) · 2016
Cited alongside, same era.
Concrete Dropout
Gal, Y., Hron, J., and Kendall, A. (2017) · 2017
Cited alongside, same era.
On Calibration of Modern Neural Networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017) · 2017
Cited alongside, same era.
Densely Connected Convolutional Networks
Huang, G., Liu, Z., van der Maaten, L., and Weinberger, K. Q. (2017) · 2017
Cited alongside, same era.
What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?
Kendall, A. and Gal, Y. (2017) · 2017
Cited alongside, same era.
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G. (2018) · 2018
Later among the works it cites.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G. (2018) · 2018
Later among the works it cites.
Fast and Scalable Bayesian Deep Learning by Weight-Perturbation in Adam
Khan, M. E., Nielsen, D., Tangkaratt, V., Lin, W., Gal, Y., and Srivastava, A. (2018) · 2018
Later among the works it cites.
Accurate Uncertainties for Deep Learning Using Calibrated Regression
Kuleshov, V., Fenner, N., and Ermon, S. (2018) · 2018
Later among the works it cites.
Visualizing the Loss Landscape of Neural Nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T. (2018) · 2018
Later among the works it cites.
Evaluating Bayesian Deep Learning Methods for Semantic Segmentation
Mukhoti, J. and Gal, Y. (2018) · 2018
Later among the works it cites.
Improving stability in deep reinforcement learning with weight averaging
Nikishin, E., Izmailov, P., Athiwaratkun, B., Podoprikhin, D., Garipov, T., Shvechikov, P., Vetrov, D., and Wilson, A. G. (2018) · 2018
Later among the works it cites.
Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
Sagun, L., Evci, U., Guney, V. U., Dauphin, Y., and Bottou, L. (2018) · 2018
Later among the works it cites.
There are many consistent explanations for unlabeled data: why you should average
Athiwaratkun, B., Finzi, M., Izmailov, P., and Wilson, A. G. (2019) · 2019
Closest in time.
Gradient descent happens in a tiny subspace
Gur-Ari, G., Roberts, D. A., and Dyer, E. (2019) · 2019
Closest in time.
Decoupled Weight Decay Regularization
Loshchilov, I. and Hutter, F. (2019) · 2019
Closest in time.
Fixing variational bayes: Deterministic variational inference for bayesian neural networks
Wu, A., Nowozin, S., Meeds, E., Turner, R. E., Hernández-Lobato, J. M., and Gaunt, A. L. (2019) · 2019
Closest in time.
Swalp: Stochastic weight averaging in low precision training
Yang, G., Zhang, T., Kirichenko, P., Bai, J., Wilson, A. G., and De Sa, C. (2019) · 2019
Closest in time.
The Unusual Effectiveness of Averaging in GAN Training
Yazici, Y., Foo, C.-S., Winkler, S., Yap, K.-H., Piliouras, G., and Chandrasekhar, V. (2019) · 2019
Closest in time.