Fetching the paper…
Reading the bibliography…
We make three related contributions motivated by the challenge of training stochastic neural networks, particularly in a PAC-Bayesian setting: (1) we show how averaging over an ensemble of stochastic neural networks enables a new class of \emph{partially-aggregated} estimators; (2) we show that these lead to provably lower-variance gradient estimates for non-differentiable signed-output networks; (3) we reformulate a PAC-Bayesian bound for these networks to derive a directly optimisable, differentiable objective and a generalisation guarantee, without using a surrogate loss or loosening the bound.
(Not) Bounding the True Error
John Langford and Rich Caruana · 1968
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Bounds for averaging classifiers
John Langford and Matthias Seeger · 2001
Earlier work this paper cites.
An improved predictive accuracy bound for averaging classifiers
Matthias Seeger, John Langford, and Nimrod Megiddo · 2001
Earlier work this paper cites.
Pac-Bayesian Supervised Classification: The Thermodynamics of Statistical Learning
Olivier Catoni · 2007
Earlier work this paper cites.
PAC-Bayesian learning of linear classifiers
Pascal Germain, Alexandre Lacasse, François Laviolette, and Mario Marchand · 2009
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Cited alongside, same era.
Weight uncertainty in neural network
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Cited alongside, same era.
Variational dropout and the local reparameterization trick
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Later among the works it cites.
Generalized Variational Inference: Three arguments for deriving new Posteriors
Jeremias Knoblauch, Jack Jewson, and Theodoros Damoulas · 2019
Later among the works it cites.
Dichotomize and generalize: PAC-Bayesian binary activated deep neural networks
Gaël Letarte, Pascal Germain, Benjamin Guedj, and Francois Laviolette · 2019
Later among the works it cites.
Monte carlo gradient estimation in machine learning
Shakir Mohamed, Mihaela Rosca, Michael Figurnov, and Andriy Mnih · 2019
Later among the works it cites.
Non-vacuous generalization bounds at the ImageNet scale: A PAC-Bayesian compression approach
Wenda Zhou, Victor Veitch, Morgane Austern, Ryan P. Adams, and Peter Orbanz · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Durk P Kingma, Tim Salimans, and Max Welling · 2015
Cited alongside, same era.
PAC-Bayesian theory meets bayesian inference
Pascal Germain, Francis Bach, Alexandre Lacoste, and Simon Lacoste-Julien · 2016
Cited alongside, same era.
Later among the works it cites.