Fetching the paper…
Reading the bibliography…
Given two networks with the same training loss on a dataset, when would they have drastically different test losses and errors? Better understanding of this question of generalization may improve practical applications of deep networks.
Efficient backprop
Yann LeCun, Leo Bottou, G. Orr, and Glaus-Robert Muller · 1999
Earlier work this paper cites.
Neural Network Learning - Theoretical Foundations
M. Anthony and P. Bartlett · 2002
Earlier work this paper cites.
Convexity, classification, and risk bounds
Peter L. Bartlett, Michael I. Jordan, and Jon D. McAuliffe · 2003
Earlier work this paper cites.
Introduction to statistical learning theory
O. Bousquet, S. Boucheron, and G. Lugosi · 2003
Earlier work this paper cites.
On the complexity of linear prediction: Risk bounds, margin bounds, and regularization
Sham M Kakade, Karthik Sridharan, and Ambuj Tewari · 2009
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Peter L. Bartlett, Dylan J. Foster, and Matus Telgarsky · 2017
Cited alongside, same era.
Size-independent sample complexity of neural networks
Noah Golowich, Alexander Rakhlin, and Ohad Shamir · 2017
Cited alongside, same era.
Fisher-rao metric, geometry, and complexity of neural networks
Tengyuan Liang, Tomaso Poggio, Alexander Rakhlin, and James Stokes · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Theory II: Landscape of the empirical risk in deep learning
T. Poggio and Q. Liao · 2017
Later among the works it cites.
Theory of deep learning III: explaining the non-overfitting puzzle
T. Poggio, Q. Liao, B. Miranda, L. Rosasco, X. Boix, J. Hidary, and H. Mhaskar · 2017
Later among the works it cites.
Theory of deep learning IIb: Optimization properties of SGD
C. Zhang, Q. Liao, A. Rakhlin, K. Sridharan, B. Miranda, N.Golowich, and T. Poggio · 2017
Later among the works it cites.
Classical generalization bounds are surprisingly tight for deep networks
Qianli Liao, Brando Miranda, Jack Hidary, and Tomaso Poggio · 2018
Closest in time.
Theory IIIb: Generalization in deep networks
T. Poggio, Q. Liao, B. Miranda, A. Banburski, X. Boix, and J. Hidary · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Closest in time.