Fetching the paper…
Reading the bibliography…
We study PAC-Bayesian generalization bounds for Multilayer Perceptrons (MLPs) with the cross entropy loss.
Stochastic relaxation, gibbs distributions, and the bayesian restoration of images
S. Geman and D. Geman · 1984
Earlier work this paper cites.
Learning representations by back-propagating errors
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1986
Earlier work this paper cites.
First and second order methods for learning: Between steepest descent and newton’s method
Roberto Battiti · 1992
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, and P. Haffner · 1998
Earlier work this paper cites.
Some pac-bayesian theorems
David A McAllester · 1999
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Geoffrey E Hinton · 2002
Earlier work this paper cites.
Pac-bayesian stochastic model selection
David A McAllester · 2003
Earlier work this paper cites.
A tutorial on energy-based learning
Yann LeCun, Sumit Chopra, Raia Hadsell, Marc’Aurelio Ranzato, and Fu Jie Huang · 2006
Earlier work this paper cites.
Tighter pac-bayes bounds
Amiran Ambroladze, Emilio Parrado-Hernández, and John S Shawe-taylor · 2007
Earlier work this paper cites.
Pac-bayesian supervised classification: the thermodynamics of statistical learning
Olivier Catoni · 2007
Earlier work this paper cites.
Graphical models, exponential families, and variational inference
Martin J Wainwright and Michael I Jordan · 2008
Earlier work this paper cites.
Pac-bayesian learning of linear classifiers
Pascal Germain, Alexandre Lacasse, François Laviolette, and Mario Marchand · 2009
Earlier work this paper cites.
Stochastic variational inference
Matthew D. Hoffman, David M. Blei, Chong Wang, and John Paisley · 2013
Cited alongside, same era.
Tighter pac-bayes bounds through distribution-dependent priors
Guy Lever, François Laviolette, and John Shawe-Taylor · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
An exact mapping between the variational renormalization group and deep learning
Pankaj Mehta and David J. Schwab · 2014
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Cited alongside, same era.
A pac-bayesian analysis of randomized learning with application to stochastic gradient descent
Ben London · 2017
Later among the works it cites.
Stochastic gradient descent as approximate bayesian inference
Stephan Mandt, Matthew D Hoffman, and David M Blei · 2017
Later among the works it cites.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro · 2017
Later among the works it cites.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, and Nathan Srebro · 2017
Later among the works it cites.
Why and when can deep-but not shallow-networks avoid the curse of dimensionality: a review
Tomaso Poggio, Hrushikesh Mhaskar, Lorenzo Rosasco, Brando Miranda, and Qianli Liao · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the properties of variational approximations of gibbs posteriors
Pierre Alquier, James Ridgway, and Nicolas Chopin · 2016
Cited alongside, same era.
Pac-bayesian theory meets bayesian inference
Pascal Germain, Francis Bach, Alexandre Lacoste, and Simon Lacoste-Julien · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Variational inference: A review for statisticians
David Blei, Alp Kucukelbir, and Jon MaAuliffe · 2017
Cited alongside, same era.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Generalization in deep learning
Kenji Kawaguchi, Leslie Pack Kaelbling, and Yoshua Bengio · 2017
Cited alongside, same era.
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Later among the works it cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Later among the works it cites.
Non-vacuous generalization bounds at the imagenet scale: a pac-bayesian compression approach
Wenda Zhou, Victor Veitch, Morgane Austern, Ryan P Adams, and Peter Orbanz · 2018
Later among the works it cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Later among the works it cites.
Dichotomize and generalize: Pac-bayesian binary activated deep neural networks
Gaël Letarte, Pascal Germain, Benjamin Guedj, and François Laviolette · 2019
Later among the works it cites.
Deterministic pac-bayesian generalization bounds for deep networks via generalizing noise-resilience
Vaishnavh Nagarajan and J Zico Kolter · 2019
Later among the works it cites.
Uniform convergence may be unable to explain generalization in deep learning
Vaishnavh Nagarajan and J Zico Kolter · 2019
Later among the works it cites.