Fetching the paper…
Reading the bibliography…
While there has been progress in developing non-vacuous generalization bounds for deep neural networks, these bounds tend to be uninformative about why deep learning works.
Deterministic pac-bayesian generalization bounds for deep networks via generalizing noise-resilience
Vaishnavh Nagarajan and J Zico Kolter · 1905
Earlier work this paper cites.
A formal theory of inductive inference. part i
Ray J Solomonoff · 1964
Earlier work this paper cites.
A Treatise of Human Nature
David Hume · 1978
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Geoffrey E Hinton and Drew Van Camp · 1993
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
No free lunch theorems for optimization
David H Wolpert and William G Macready · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
PAC-Bayesian Model Averaging
David A McAllester · 1999
Earlier work this paper cites.
(not) bounding the true error
John Langford and Rich Caruana · 2001
Earlier work this paper cites.
Bounds for Averaging Classifiers
John Langford and Matthias Seeger · 2001
Earlier work this paper cites.
Quantitatively tight sample complexity bounds
John Langford · 2002
Earlier work this paper cites.
Information theory, inference and learning algorithms
David JC MacKay, David JC Mac Kay, et al · 2003
Earlier work this paper cites.
A Note on the PAC Bayesian Theorem
Andreas Maurer · 2004
Earlier work this paper cites.
Toward a justification of meta-learning: Is the no free lunch theorem a show-stopper
Christophe Giraud-Carrier and Foster Provost · 2005
Earlier work this paper cites.
Very sparse random projections
Ping Li, Trevor J. Hastie, and Kenneth Ward Church · 2006
Earlier work this paper cites.
PAC-Bayesian Supervised Classification: the Thermodynamics of Statistical Learning
Olivier Catoni · 2007
Earlier work this paper cites.
An empirical evaluation of deep architectures on problems with many factors of variation
Hugo Larochelle, Dumitru Erhan, Aaron C. Courville, James Bergstra, and Yoshua Bengio · 2007
Earlier work this paper cites.
Algorithmic complexity, 2008
Marcus Hutter et al · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
Alex Krizhevsky · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning, 2011
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Estimating or Propagating Gradients through Stochastic Neurons for Conditional Computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Fastfood: Approximate kernel expansions in loglinear time
Quoc V. Le, Tamás Sarlós, and Alex Smola · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N. Dauphin, Razvan Pascanu, Çaglar Gülçehre, KyungHyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
Weight uncertainty in neural network
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Ian J. Goodfellow and Oriol Vinyals · 2015
Cited alongside, same era.
Group equivariant convolutional networks
Taco Cohen and Max Welling · 2016
Cited alongside, same era.
Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Song Han, Huizi Mao, and William J. Dally · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Binarized neural networks
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Cited alongside, same era.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V. Le · 2019
Later among the works it cites.
General e(2)-equivariant steerable cnns
Maurice Weiler and Gabriele Cesa · 2019
Later among the works it cites.
Understanding straight-through estimator in training activation quantized neural nets
Penghang Yin, Jiancheng Lyu, Shuai Zhang, Stanley J. Osher, Yingyong Qi, and Jack Xin · 2019
Later among the works it cites.
Non-vacuous generalization bounds at the imagenet scale: a pac-bayesian compression approach
Wenda Zhou, Victor Veitch, Morgane Austern, Ryan P. Adams, and Peter Orbanz · 2019
Later among the works it cites.
Experiment tracking with weights and biases, 2020
Lukas Biewald · 2020
Later among the works it cites.
Generalizing convolutional neural networks for equivariance to lie groups on arbitrary continuous data
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer T. Chayes, Levent Sagun, and Riccardo Zecchina · 2017
Cited alongside, same era.
A survey of model compression and acceleration for deep neural networks
Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang · 2017
Cited alongside, same era.
Towards the limit of network quantization
Yoojin Choi, Mostafa El-Khamy, and Jungwon Lee · 2017
Cited alongside, same era.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Gintare Karolina Dziugaite and Daniel M. Roy · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Cited alongside, same era.
A Strongly Quasiconvex PAC-Bayesian Bound
Niklas Thiemann, Christian Igel, Oliver Wintenberger, and Yevgeny Seldin · 2017
Cited alongside, same era.
Marc Finzi, Samuel Stanton, Pavel Izmailov, and Andrew Gordon Wilson · 2020
Later among the works it cites.
On the benefits of invariance in neural networks
Clare Lyle, Mark van der Wilk, Marta Kwiatkowska, Yarin Gal, and Benjamin Bloem-Reddy · 2020
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2020
Later among the works it cites.
In defense of uniform convergence: Generalization via derandomization with an application to interpolating predictors
Jeffrey Negrea, Gintare Karolina Dziugaite, and Daniel Roy · 2020
Later among the works it cites.
On the generalization benefit of noise in stochastic gradient descent
Samuel L. Smith, Erich Elsen, and Soham De · 2020
Later among the works it cites.
On the noisy gradient descent that generalizes as SGD
Jingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan, Vladimir Braverman, and Zhanxing Zhu · 2020
Later among the works it cites.
On the sample complexity of learning under geometric stability
Alberto Bietti, Luca Venturi, and Joan Bruna · 2021
Later among the works it cites.
On the Role of Data in Pac-Bayes Bounds
Gintare Karolina Dziugaite, Kyle Hsu, Waseem Gharbieh, Gabriel Aprino, and Daniel M. Roy · 2021
Later among the works it cites.
Provably strict generalisation benefit for equivariant models
Bryn Elesedy and Sheheryar Zaidi · 2021
Later among the works it cites.
What are bayesian neural network posteriors really like?
Pavel Izmailov, Sharad Vikram, Matthew D Hoffman, and Andrew Gordon Gordon Wilson · 2021
Later among the works it cites.
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al · 2021
Later among the works it cites.
Partial transfusion: on the expressive influence of trainable batch norm parameters for transfer learning
Fahdi Kanavati and Masayuki Tsuneki · 2021
Later among the works it cites.
On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs)
Zhiyuan Li, Sadhika Malladi, and Sanjeev Arora · 2021
Later among the works it cites.
Tighter Risk Certificates for Neural Networks
María Pérez-Ortiz, Omar Rivasplata, John Shawe-Taylor, and Csaba Szepersvári · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Later among the works it cites.
Understanding the generalization benefit of model invariance from a data perspective
Sicheng Zhu, Bang An, and Furong Huang · 2021
Later among the works it cites.
Nan Ding, Xi Chen, Tomer Levinboim, Beer Changpinyo, and Radu Soricut · 2022
Closest in time.
Group symmetry in pac learning
Bryn Elesedy · 2022
Closest in time.
Stochastic Training Is Not Necessary For Generalization
Jonas Geiping, Micah Goldblum, Phillip E. Pope, Michael Moeller, and Tom Goldstein · 2022
Closest in time.