Fetching the paper…
Reading the bibliography…
With a goal of understanding what drives generalization in deep networks, we consider several recently suggested explanations, including norm-based control, sharpness and robustness.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
P. L. Bartlett · 1998
Earlier work this paper cites.
Almost linear vc dimension bounds for piecewise polynomial networks
P. L. Bartlett, V. Maiorov, and R. Meir · 1998
Earlier work this paper cites.
Some PAC-Bayesian theorems
D. A. McAllester · 1998
Earlier work this paper cites.
The connection between regularization operators and support vector kernels
A. J. Smola, B. Schölkopf, and K.-R. Müller · 1998
Earlier work this paper cites.
PAC-Bayesian model averaging
D. A. McAllester · 1999
Earlier work this paper cites.
Regularization networks and support vector machines
T. Evgeniou, M. Pontil, and T. Poggio · 2000
Earlier work this paper cites.
(not) bounding the true error
J. Langford and R. Caruana · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
P. L. Bartlett and S. Mendelson · 2002
Earlier work this paper cites.
Simplified pac-bayesian margin bounds
D. McAllester · 2003
Cited alongside, same era.
Distance-based classification with lipschitz functions
U. v. Luxburg and O. Bousquet · 2004
Cited alongside, same era.
Rank, trace-norm and max-norm
N. Srebro and A. Shraibman · 2005
Cited alongside, same era.
Maximum-margin matrix factorization
N. Srebro, J. Rennie, and T. S. Jaakkola · 2005
Cited alongside, same era.
Neural network learning: Theoretical foundations
M. Anthony and P. L. Bartlett · 2009
Cited alongside, same era.
Robustness and generalization
H. Xu and S. Mannor · 2012
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Entropy-sgd: Biasing gradient descent into wide valleys
P. Chaudhari, A. Choromanska, S. Soatto, and Y. LeCun · 2016
Later among the works it cites.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2016
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Later among the works it cites.
Generalization error of invariant classifiers
J. Sokolic, R. Giryes, G. Sapiro, and M. R. Rodrigues · 2016
Later among the works it cites.
The impact of the nonlinearity on the VC-dimension of a deep network
P. L. Bartlett · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Path-SGD: Path-normalized optimization in deep neural networks
B. Neyshabur, R. Salakhutdinov, and N. Srebro
Cited in the paper.
Norm-based capacity control in neural networks
B. Neyshabur, R. Tomioka, and N. Srebro
Cited in the paper.
In search of the real inductive bias: On the role of implicit regularization in deep learning
B. Neyshabur, R. Tomioka, and N. Srebro
Cited in the paper.
Data-dependent path normalization in neural networks
B. Neyshabur, R. Tomioka, R. Salakhutdinov, and N. Srebro
Cited in the paper.
G. K. Dziugaite and D. M. Roy · 2017
Closest in time.
Nearly-tight vc-dimension bounds for piecewise linear neural networks
N. Harvey, C. Liaw, and A. Mehrabian · 2017
Closest in time.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Closest in time.