Fetching the paper…
Reading the bibliography…
Existing generalization bounds fail to explain crucial factors that drive the generalization of modern neural networks.
On the uniform convergence of relative frequencies of events to their probabilities
V. N. Vapnik and A. Y. Chervonenkis · 1971
Earlier work this paper cites.
On the method of bounded differences
C. McDiarmid et al · 1989
Earlier work this paper cites.
A simple weight decay can improve generalization
A. Krogh and J. Hertz · 1991
Earlier work this paper cites.
Scale-sensitive dimensions, uniform convergence, and learnability
N. Alon, S. Ben-David, N. Cesa-Bianchi, and D. Haussler · 1997
Earlier work this paper cites.
A PAC analysis of a bayesian estimator
J. Shawe-Taylor and R. C. Williamson · 1997
Earlier work this paper cites.
Almost linear vc dimension bounds for piecewise polynomial networks
P. Bartlett, V. Maiorov, and R. Meir · 1998
Earlier work this paper cites.
Some pac-bayesian theorems
D. A. McAllester · 1998
Earlier work this paper cites.
Rademacher penalties and structural risk minimization
V. Koltchinskii · 2001
Earlier work this paper cites.
(Not) bounding the true error
J. Langford and R. Caruana · 2001
Earlier work this paper cites.
Data-dependent margin-based generalization bounds for classification
A. Antos, B. Kégl, T. Linder, and G. Lugosi · 2002
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
P. L. Bartlett and S. Mendelson · 2002
Earlier work this paper cites.
Real analysis and probability , volume 74 of Cambridge Studies in Advanced Mathematics
R. M. Dudley · 2002
Earlier work this paper cites.
Local Rademacher complexities
P. L. Bartlett, O. Bousquet, and S. Mendelson · 2005
Earlier work this paper cites.
PAC-Bayesian supervised classification: the thermodynamics of statistical learning
O. Catoni · 2007
Earlier work this paper cites.
Optimal transport: old and new
C. Villani · 2009
Earlier work this paper cites.
Sample complexity of testing the manifold hypothesis
H. Narayanan and S. Mitter · 2010
Earlier work this paper cites.
Dimensions, embeddings, and attractors
J. C. Robinson · 2011
Earlier work this paper cites.
Manifold estimation and singular deconvolution under Hausdorff loss
C. R. Genovese, M. Perone-Pacifico, I. Verdinelli, and L. Wasserman · 2012
Earlier work this paper cites.
On the mean speed of convergence of empirical and occupation measures in Wasserstein distance
E. Boissard and T. Le Gouic · 2014
Earlier work this paper cites.
On the number of linear regions of deep neural networks
G. F. Montufar, R. Pascanu, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Norm-based capacity control in neural networks
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Earlier work this paper cites.
On the properties of variational approximations of Gibbs posteriors
P. Alquier, J. Ridgway, N. Chopin, and Y. W. Teh · 2016
Earlier work this paper cites.
Testing the manifold hypothesis
C. Fefferman, S. Mitter, and H. Narayanan · 2016
Earlier work this paper cites.
On the depth of deep neural networks: A theoretical view
S. Sun, W. Chen, L. Wang, X. Liu, and T.-Y. Liu · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
P. L. Bartlett, D. J. Foster, and M. J. Telgarsky · 2017
Earlier work this paper cites.
Parseval networks: Improving robustness to adversarial examples
M. Cisse, P. Bojanowski, E. Grave, Y. Dauphin, and N. Usunier · 2017
Earlier work this paper cites.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
G. K. Dziugaite and D. M. Roy · 2017
Earlier work this paper cites.
Efficient regression in metric spaces via approximate lipschitz extension
L.-A. Gottlieb, A. Kontorovich, and R. Krauthgamer · 2017
Earlier work this paper cites.
Nearly-tight VC-dimension bounds for piecewise linear neural networks
N. Harvey, C. Liaw, and A. Mehrabian · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Cited alongside, same era.
Exploring generalization in deep learning
B. Neyshabur, S. Bhojanapalli, D. Mcallester, and N. Srebro · 2017
Cited alongside, same era.
Robust large margin deep neural networks
J. Sokolic, R. Giryes, G. Sapiro, and M. R. D. Rodrigues · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
S. Arora, R. Ge, B. Neyshabur, and Y. Zhang · 2018
Cited alongside, same era.
Neural ordinary differential equations
R. T. Chen, Y. Rubanova, J. Bettencourt, and D. K. Duvenaud · 2018
Cited alongside, same era.
High-dimensional statistics: A non-asymptotic viewpoint
M. J. Wainwright · 2019
Later among the works it cites.
Sharp asymptotic and finite-sample rates of convergence of empirical measures in Wasserstein distance
J. Weed and F. Bach · 2019
Later among the works it cites.
A short note on learning discrete distributions
C. L. Canonne · 2020
Later among the works it cites.
A generative adversarial network approach to calibration of local stochastic volatility models
C. Cuchiero, W. Khosrawi, and J. Teichmann · 2020
Later among the works it cites.
Size-independent sample complexity of neural networks
N. Golowich, A. Rakhlin, and O. Shamir · 2020
Later among the works it cites.
Estimating full lipschitz constants of deep neural networks
C. Herrera, F. Krach, and J. Teichmann · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Data-dependent PAC-Bayes priors via differential privacy
G. K. Dziugaite and D. M. Roy · 2018
Cited alongside, same era.
Lipschitz regularized deep neural networks generalize and are adversarially robust
C. Finlay, J. Calder, B. Abbasi, and A. Oberman · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
S. Gunasekar, J. D. Lee, D. Soudry, and N. Srebro · 2018
Cited alongside, same era.
Minimax statistical learning with wasserstein distances
J. Lee and M. Raginsky · 2018
Cited alongside, same era.
Data-driven distributionally robust optimization using the wasserstein metric: Performance guarantees and tractable reformulations
P. Mohajerin Esfahani and D. Kuhn · 2018
Cited alongside, same era.
A PAC-Bayesian approach to spectrally-normalized margin bounds for neural networks
B. Neyshabur, S. Bhojanapalli, and N. Srebro · 2018
Cited alongside, same era.
Exactly computing the local lipschitz constant of relu networks
M. Jordan and A. G. Dimakis · 2020
Later among the works it cites.
Empirical measures: regularity is a counter-curse to dimensionality
B. R. Kloeckner · 2020
Later among the works it cites.
Wasserstein smoothing: Certified robustness against wasserstein adversarial attacks
A. Levine and S. Feizi · 2020
Later among the works it cites.
Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks
M. Li, M. Soltanolkotabi, and S. Oymak · 2020
Later among the works it cites.
Grokking deep reinforcement learning
M. Morales · 2020
Later among the works it cites.
Distributionally robust neural networks
S. Sagawa, P. W. Koh, T. B. Hashimoto, and P. Liang · 2020
Later among the works it cites.
The phase diagram of approximation rates for deep neural networks
D. Yarotsky and A. Zhevnerchuk · 2020
Later among the works it cites.
Failures of model-dependent generalization bounds for least-norm interpolation
P. L. Bartlett and P. M. Long · 2021
Later among the works it cites.
Intrinsic dimension estimation
A. Block, Z. Jia, Y. Polyanskiy, and A. Rakhlin · 2021
Later among the works it cites.
Generalised lipschitz regularisation equals distributional robustness
Z. Cranko, Z. Shi, X. Zhang, R. Nock, and S. Kornblith · 2021
Later among the works it cites.
On the approximation of functions by tanh neural networks
T. De Ryck, S. Lanthaler, and S. Mishra · 2021
Later among the works it cites.
Regularisation of neural networks by enforcing lipschitz continuity
H. Gouk, E. Frank, B. Pfahringer, and M. J. Cree · 2021
Later among the works it cites.
Approximation rates for neural networks with encodable weights in smoothness spaces
I. Gühring and M. Raslan · 2021
Later among the works it cites.
NEU: A meta-algorithm for universal uap-invariant feature representation
C. Hyndman and A. Kratsios · 2021
Later among the works it cites.
Robustness to pruning predicts generalization in deep neural networks
L. Kuhn, C. Lyle, A. N. Gomez, J. Rothfuss, and Y. Gal · 2021
Later among the works it cites.
An alternative probabilistic interpretation of the Huber loss
G. P. Meyer · 2021
Later among the works it cites.
Neural rough differential equations for long time series
J. Morrill, C. Salvi, P. Kidger, and J. Foster · 2021
Later among the works it cites.
The intrinsic dimension of images and its impact on learning
P. Pope, C. Zhu, A. Abdelkader, M. Goldblum, and T. Goldstein · 2021
Later among the works it cites.
On the generalization properties of adversarial training
Y. Xing, Q. Song, and G. Cheng · 2021
Later among the works it cites.
Finite-sample guarantees for wasserstein distributionally robust optimization: Breaking the curse of dimensionality
R. Gao · 2022
Closest in time.
Distributionally robust stochastic optimization with wasserstein distance
R. Gao and A. Kleywegt · 2022
Closest in time.
Adversarial robustness of sparse local lipschitz predictors
R. Muthukumar and J. Sulam · 2023
Closest in time.