Fetching the paper…
Reading the bibliography…
Two main obstacles preventing the widespread adoption of variational Bayesian neural networks are the high parameter overhead that makes them infeasible on large networks, and the difficulty of implementation, which can be thought of as "programming overhead." MC dropout [Gal and Ghahramani, 2016] is popular because it sidesteps these obstacles.
A practical Bayesian framework for backpropagation networks
David JC MacKay · 1992
Earlier work this paper cites.
Bayesian training of backpropagation networks by the hybrid Monte Carlo method
Radford M Neal · 1992
Earlier work this paper cites.
Approximating posterior distributions in belief networks using mixtures
Christopher M Bishop, Neil D Lawrence, Tommi Jaakkola, and Michael I Jordan · 1998
Earlier work this paper cites.
Improving the mean field approximation via the use of mixture distributions
Tommi S Jaakkola and Michael I Jordan · 1998
Earlier work this paper cites.
Bayesian approach for neural networks—review and case studies
Jouko Lampinen and Aki Vehtari · 2001
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Sparse spectrum gaussian process regression
Miguel Lázaro-Gredilla, Joaquin Quiñonero Candela, Carl Edward Rasmussen, and Aníbal R. Figueiras-Vidal · 2010
Earlier work this paper cites.
Practical variational inference for neural networks
Alex Graves · 2011
Earlier work this paper cites.
Machine learning: a probabilistic perspective
Kevin P Murphy · 2012
Earlier work this paper cites.
Nonparametric variational inference
Samuel Gershman, Matt Hoffman, and David Blei · 2012
Earlier work this paper cites.
Regularization of neural networks using Dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann Le Cun, and Rob Fergus · 2013
Earlier work this paper cites.
Philosophy and the practice of bayesian statistics
Andrew Gelman and Cosma Rohilla Shalizi · 2013
Earlier work this paper cites.
Understanding predictive information criteria for Bayesian models
Andrew Gelman, Jessica Hwang, and Aki Vehtari · 2014
Earlier work this paper cites.
Variational Bayesian inference with Gaussian-mixture approximations
O. Zobay · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Probabilistic backpropagation for scalable learning of Bayesian neural networks
José Miguel Hernández-Lobato and Ryan Adams · 2015
Earlier work this paper cites.
Weight uncertainty in neural networks
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Durk P Kingma, Tim Salimans, and Max Welling · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Cited alongside, same era.
Dropout as a Bayesian approximation: Appendix
Yarin Gal and Zoubin Ghahramani · 2015
Cited alongside, same era.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
Deepfool: a simple and accurate method to fool deep neural networks
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, and Pascal Frossard · 2016
Cited alongside, same era.
Dynamic layer normalization for adaptive neural acoustic modeling in speech recognition
Taesup Kim, Inchul Song, and Yoshua Bengio · 2017
Later among the works it cites.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville · 2017
Later among the works it cites.
Variational boosting: Iteratively refining posterior approximations
Andrew C Miller, Nicholas J Foti, and Ryan P Adams · 2017
Later among the works it cites.
Variational particle approximations
Ardavan Saeedi, Tejas D Kulkarni, Vikash K Mansinghka, and Samuel J Gershman · 2017
Later among the works it cites.
Snapshot ensembles: Train 1, get m for free
Gao Huang, Yixuan Li, Geoff Pleiss, Zhuang Liu, John E Hopcroft, and Kilian Q Weinberger · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Structured and efficient variational deep learning with matrix Gaussian posteriors
Christos Louizos and Max Welling · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Hierarchical variational models
Rajesh Ranganath, Dustin Tran, and David Blei · 2016
Cited alongside, same era.
Stochastic backpropagation through mixture density distributions
Alex Graves · 2016
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Later among the works it cites.
What uncertainties do we need in bayesian deep learning for computer vision?
Alex Kendall and Yarin Gal · 2017
Later among the works it cites.
Automatic differentiation variational inference
Alp Kucukelbir, Dustin Tran, Rajesh Ranganath, Andrew Gelman, and David M Blei · 2017
Later among the works it cites.
Understanding the disharmony between dropout and batch normalization by variance shift
Xiang Li, Shuo Chen, Xiaolin Hu, and Jian Yang · 2018
Later among the works it cites.
Tim Pearce, Nicolas Anastassacos, Mohamed Zaki, and Andy Neely · 2018
Later among the works it cites.
K for the price of 1: Parameter efficient multi-task and transfer learning
Pramod Kaushik Mudrakarta, Mark Sandler, Andrey Zhmoginov, and Andrew Howard · 2018
Later among the works it cites.
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila · 2018
Later among the works it cites.
Feature-wise transformations
Vincent Dumoulin, Ethan Perez, Nathan Schucher, Florian Strub, Harm de Vries, Aaron Courville, and Yoshua Bengio · 2018
Later among the works it cites.
Adversarial distillation of Bayesian neural network posteriors
Kuan-Chieh Wang, Paul Vicol, James Lucas, Li Gu, Roger Grosse, and Richard Zemel · 2018
Later among the works it cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Later among the works it cites.
An empirical model of large-batch training
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team · 2018
Later among the works it cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2019
Closest in time.
Do Imagenet classifiers generalize to Imagenet?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2019
Closest in time.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Closest in time.
GPU specs database
TechPowerUp · 2019
Closest in time.