Fetching the paper…
Reading the bibliography…
Modern neural networks are highly overparameterized, with capacity to substantially overfit to training data.
Stochastic complexity and modeling
Jorma Rissanen · 1986
Earlier work this paper cites.
Occam’s razor
Anselm Blumer, Andrzej Ehrenfeucht, David Haussler, and Manfred K. Warmuth · 1987
Earlier work this paper cites.
Bayesian model comparison and backprop nets
David MacKay · 1992
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Geoffrey E. Hinton and Drew van Camp · 1993
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Some PAC-Bayesian theorems
David A. McAllester · 1999
Earlier work this paper cites.
Occam’s razor
Carl Edward Rasmussen and Zoubin Ghahramani · 2001
Earlier work this paper cites.
Quantitatively tight sample complexity bounds
John Langford · 2002
Earlier work this paper cites.
(not) bounding the true error
John Langford and Rich Caruana · 2002
Earlier work this paper cites.
Simplified pac-bayesian margin bounds
David McAllester · 2003
Earlier work this paper cites.
Pac-Bayesian supervised classification: the thermodynamics of statistical learning , volume 56 of Institute of Mathematical Statistics Lecture Notes—Monograph Series
Olivier Catoni · 2007
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
A PAC-Bayesian tutorial with A dropout bound, 2013
David A. McAllester · 2013
Cited alongside, same era.
PAC-Bayesian Theory for Transductive Learning
Luc Bégin, Pascal Germain, François Laviolette, and Jean-Francis Roy · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonoyan and Andrew Zisserman · 2014
Cited alongside, same era.
Subdominant dense clusters allow for simple learning and high computational performance in neural networks with discrete synapses
Carlo Baldassi, Alessandro Ingrosso, Carlo Lucibello, Luca Saglietti, and Riccardo Zecchina · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Cited alongside, same era.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Gintare Karolina Dziugaite and Daniel M. Roy · 2017
Later among the works it cites.
Nearly-tight VC-dimension bounds for piecewise linear neural networks
Nick Harvey, Christopher Liaw, and Abbas Mehrabian · 2017
Later among the works it cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications, 2017
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Later among the works it cites.
A tutorial on pac-bayesian theory
François Laviolette · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Carlo Baldassi, Christian Borgs, Jennifer T. Chayes, Alessandro Ingrosso, Carlo Lucibello, Luca Saglietti, and Riccardo Zecchina · 2016
Cited alongside, same era.
Dynamic network surgery for efficient DNNs
Yiwen Guo, Anbang Yao, and Yurong Chen · 2016
Cited alongside, same era.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Song Han, Huizi Mao, and William J Dally · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqin Ren, and Jian Sun · 2016
Cited alongside, same era.
SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and < 0.5mb model size
Forrest N Iandola, Song Han, Matthew W Moskewicz, Khalid Ashraf, William J Dally, and Kurt Keutzer · 2016
Cited alongside, same era.
Generalization Theory and Deep Nets, An introduction
Sanjeev Arora · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks, 2017
Peter Bartlett, Dylan J. Foster, and Matus Telgarsky · 2017
Cited alongside, same era.
Behnam Neyshabur, Srinadh Bhojanapalli, David Mcallester, and Nati Srebro · 2017
Later among the works it cites.
Nonparametric regression using deep neural networks with relu activation function, August 2017
J. Schmidt-Hieber · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
To prune, or not to prune: exploring the efficacy of model pruning for model compression
Michael Zhu and Suyog Gupta · 2017
Later among the works it cites.
Stronger generalization bounds for deep nets via a compression approach, February 2018
S. Arora, R. Ge, B. Neyshabur, and Y. Zhang · 2018
Closest in time.
Model compression and acceleration for deep neural networks: The principles, progress, and challenges
Y. Cheng, D. Wang, P. Zhou, and T. Zhang · 2018
Closest in time.
Entropy-sgd optimizes the prior of a pac-bayes bound: Generalization properties of entropy-sgd and data-dependent priors
Gintare Karolina Dziugaite and Daniel M. Roy · 2018
Closest in time.
A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, and Nathan Srebro · 2018
Closest in time.