Fetching the paper…
Reading the bibliography…
Generalization of deep networks has been of great interest in recent years, resulting in a number of theoretically and empirically motivated complexity measures.
Generalization in deep networks: The role of distance from initialization
Nagarajan, V. and Kolter, J. Z. (2019b) · 1901
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V. (2019) · 1902
Earlier work this paper cites.
Size-free generalization bounds for convolutional neural networks
Long, P. M. and Sedghi, H. (2019) · 1905
Earlier work this paper cites.
Deterministic pac-bayesian generalization bounds for deep networks via generalizing noise-resilience
Nagarajan, V. and Kolter, J. Z. (2019a) · 1905
Earlier work this paper cites.
Data-dependent sample complexity of deep neural networks via lipschitz augmentation
Wei, C. and Ma, T. (2019a) · 1905
Earlier work this paper cites.
A new measure of rank correlation
Kendall, M. G. (1938) · 1938
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
Vapnik, V. N. and Chervonenkis, A. Y. (1971) · 1971
Earlier work this paper cites.
Equivalence and synthesis of causal models
Verma, T. and Pearl, J. (1991) · 1991
Earlier work this paper cites.
Pac-bayesian model averaging
McAllester, D. A. (1999) · 1999
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S. (2002) · 2002
Earlier work this paper cites.
Robustness of a network of networks
Gao, J., Buldyrev, S. V., Havlin, S., and Stanley, H. E. (2011) · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. (2011) · 2011
Earlier work this paper cites.
Foundations of machine learning. adaptive computation and machine learning
Mohri, M., Rostamizadeh, A., and Talwalkar, A. (2012) · 2012
Earlier work this paper cites.
The cifar-10 dataset
Krizhevsky, A., Nair, V., and Hinton, G. (2014) · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N. (2014) · 2014
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y. (2015) · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. (2016) · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Regularizing and optimizing lstm language models
Merity, S., Keskar, N. S., and Socher, R. (2017) · 2017
Later among the works it cites.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N. (2017) · 2017
Later among the works it cites.
Regularizing neural networks by penalizing confident output distributions
Pereyra, G., Tucker, G., Chorowski, J., Kaiser, Ł., and Hinton, G. (2017) · 2017
Later among the works it cites.
Pac-bayesian margin bounds for convolutional neural networks
Pitas, K., Davies, M., and Vandergheynst, P. (2017) · 2017
Later among the works it cites.
A bayesian perspective on generalization and stochastic gradient descent
Smith, S. L. and Le, Q. V. (2017) · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P. (2016) · 2016
Cited alongside, same era.
Dudley-pollard packing theorem
Kontorovich, A. (2016) · 2016
Cited alongside, same era.
Fantastic beasts and where to find them
Rowling, J. K. (2016) · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., and Telgarsky, M. J. (2017) · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y. (2017) · 2017
Cited alongside, same era.
Dziugaite, G. K. and Roy, D. M. (2017) · 2017
Cited alongside, same era.
Size-independent sample complexity of neural networks
Golowich, N., Rakhlin, A., and Shamir, O. (2017) · 2017
Cited alongside, same era.
Later among the works it cites.
Stronger generalization bounds for deep nets via a compression approach
Arora, S., Ge, R., Neyshabur, B., and Zhang, Y. (2018) · 2018
Later among the works it cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Chaudhari, P. and Soatto, S. (2018) · 2018
Later among the works it cites.
Large margin deep networks for classification
Elsayed, G., Krishnan, D., Mobahi, H., Regan, K., and Bengio, S. (2018) · 2018
Later among the works it cites.
Predicting the generalization gap in deep networks with margin distributions
Jiang, Y., Krishnan, D., Mobahi, H., and Bengio, S. (2018) · 2018
Later among the works it cites.
Sensitivity and generalization in neural networks: an empirical study
Novak, R., Bahri, Y., Abolafia, D. A., Pennington, J., and Sohl-Dickstein, J. (2018) · 2018
Later among the works it cites.
The singular values of convolutional layers
Sedghi, H., Gupta, V., and Long, P. M. (2018) · 2018
Later among the works it cites.
Nearly-tight vc-dimension and pseudodimension bounds for piecewise linear neural networks
Bartlett, P. L., Harvey, N., Liaw, C., and Mehrabian, A. (2019) · 2019
Closest in time.
Over-parametrization in deep rl and causal graphs for deep learning theory
Neal, B. (2019) · 2019
Closest in time.