Fetching the paper…
Reading the bibliography…
This paper focuses on understanding how the generalization error scales with the amount of the training data for deep neural networks (DNNs).
On the density of families of sets
Norbert Sauer · 1972
Earlier work this paper cites.
Estimation of Dependences Based on Empirical Data: Springer Series in Statistics (Springer Series in Statistics)
Vladimir Vapnik · 1982
Earlier work this paper cites.
A general lower bound on the number of examples needed for learning
Andrzej Ehrenfeucht, David Haussler, Michael Kearns, and Leslie Valiant · 1989
Earlier work this paper cites.
Learning curves: Asymptotic values and rate of convergence
Corinna Cortes, Lawrence D Jackel, Sara A Solla, Vladimir Vapnik, and John S Denker · 1994
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1994
Earlier work this paper cites.
Limits on learning machine accuracy imposed by data quality
Corinna Cortes, Lawrence D Jackel, and Wan-Ping Chiang · 1995
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik · 1995
Earlier work this paper cites.
Scale-sensitive dimensions, uniform convergence, and learnability
Noga Alon, Shai Ben-David, Nicolo Cesa-Bianchi, and David Haussler · 1997
Earlier work this paper cites.
Monte carlo and quasi-monte carlo methods
Russel E. Caflisch · 1998
Earlier work this paper cites.
Statistical Learning Theory
Vladimir N. Vapnik · 1998
Earlier work this paper cites.
Neural Network Learning: Theoretical Foundations
Martin Anthony and Peter L. Bartlett · 1999
Earlier work this paper cites.
Modeling decision tree performance with the power law
Lewis J. Frey and Douglas H. Fisher · 1999
Earlier work this paper cites.
Pac-bayesian model averaging
David A McAllester · 1999
Earlier work this paper cites.
Upper and lower bounds on the learning curve for gaussian processes
Christopher KI Williams and Francesco Vivarelli · 2000
Earlier work this paper cites.
Modelling classification performance for large data sets
Baohua Gu, Feifang Hu, and Huan Liu · 2001
Earlier work this paper cites.
Rademacher penalties and structural risk minimization
Vladimir Koltchinskii · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Gaussian process regression with mismatched models
Peter Sollich · 2002
Earlier work this paper cites.
Covering number bounds of certain regularized linear function classes
Tong Zhang · 2002
Earlier work this paper cites.
Introduction to statistical learning theory
Olivier Bousquet, Stéphane Boucheron, and Gábor Lugosi · 2003
Earlier work this paper cites.
Pac-bayes & margins
John Langford and John Shawe-Taylor · 2003
Earlier work this paper cites.
Simplified pac-bayesian margin bounds
David McAllester · 2003
Earlier work this paper cites.
Distance-based classification with lipschitz functions
Ulrike von Luxburg and Olivier Bousquet · 2004
Cited alongside, same era.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Cited alongside, same era.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Cited alongside, same era.
Statistical learning theory: a tutorial
Sanjeev R. Kulkarni and Gilbert Harman · 2011
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Cited alongside, same era.
Pruning convolutional neural networks for resource efficient inference
Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz · 2017
Later among the works it cites.
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Later among the works it cites.
The AI Trinity: Data + Algorithms + Infrastructure
Anima Anandkumar · 2018
Later among the works it cites.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Later among the works it cites.
Statistical learning theory
Rui Castro · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The nature of statistical learning theory
Vladimir Vapnik · 2013
Cited alongside, same era.
Exploiting linear structure within convolutional networks for efficient evaluation
Emily L Denton, Wojciech Zaremba, Joan Bruna, Yann LeCun, and Rob Fergus · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Cited alongside, same era.
State-of-the-art speech recognition with sequence-to-sequence models
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Ekaterina Gonina, et al · 2018
Later among the works it cites.
Low rank structure of learned representations
Amartya Sanyal, Varun Kanade, and Philip H. S. Torr · 2018
Later among the works it cites.
To prune, or not to prune: Exploring the efficacy of pruning for model compression
Michael Zhu and Suyog Gupta · 2018
Later among the works it cites.
Learning representations for neural network-based classification using the information bottleneck principle
Rana Ali Amjad and Bernhard Claus Geiger · 2019
Later among the works it cites.
Intrinsic dimension of data representations in deep neural networks
Alessio Ansuini, Alessandro Laio, Jakob H Macke, and Davide Zoccolan · 2019
Later among the works it cites.
Nearly-tight vc-dimension and pseudodimension bounds for piecewise linear neural networks
Peter L. Bartlett, Nick Harvey, Christopher Liaw, and Abbas Mehrabian · 2019
Later among the works it cites.
Learning curves for deep neural networks: a gaussian field theory perspective
Omry Cohen, Or Malka, and Zohar Ringel · 2019
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Estimating information flow in deep neural networks
Ziv Goldfeld, Ewout Van Den Berg, Kristjan Greenewald, Igor Melnyk, Nam Nguyen, Brian Kingsbury, and Yury Polyanskiy · 2019
Later among the works it cites.
The role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2019
Later among the works it cites.
Using effective dimension to analyze feature transformations in deep neural networks
Kavya Ravichandran, Ajay Jain, and Alexander Rakhlin · 2019
Later among the works it cites.
Dimensionality compression and expansion in deep neural networks
Stefano Recanatesi, Matthew Farrell, Madhu Advani, Timothy Moore, Guillaume Lajoie, and Eric Shea-Brown · 2019
Later among the works it cites.
Asymptotic learning curves of kernel methods: empirical data vs teacher-student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2019
Later among the works it cites.
Non-vacuous generalization bounds at the imagenet scale: a PAC-bayesian compression approach
Wenda Zhou, Victor Veitch, Morgane Austern, Ryan P. Adams, and Peter Orbanz · 2019
Later among the works it cites.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Dilip Krishnan, Hossein Mobahi, and Samy Bengio · 2020
Later among the works it cites.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Dilip Krishnan, Hossein Mobahi, and Samy Bengio · 2020
Later among the works it cites.