Fetching the paper…
Reading the bibliography…
Modern neural network architectures often generalize well despite containing many more parameters than the size of the training dataset.
Bemerkungen zur theorie der beschränkten bilinearformen mit unendlich vielen veränderlichen
Schur, J · 1911
Earlier work this paper cites.
Perturbation bounds for matrix square roots and pythagorean sums
Schmitt, B. A · 1992
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Bartlett, P. L · 1998
Earlier work this paper cites.
Almost linear vc dimension bounds for piecewise polynomial networks
Bartlett, P. L., Maiorov, V., and Meir, R · 1999
Earlier work this paper cites.
The concentration of measure phenomenon
Ledoux, M · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. L., and Mendelson, S · 2002
Earlier work this paper cites.
Neural network learning: Theoretical foundations
Anthony, M., and Bartlett, P. L · 2009
Earlier work this paper cites.
A useful variant of the davis–kahan theorem for statisticians
Yu, Y., Wang, T., and Samworth, R. J · 2014
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2015
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
A vector-contraction inequality for rademacher complexities
Maurer, A · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Earlier work this paper cites.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A., Globerson, A., Malach, E., and Shalev-Shwartz, S · 2017
Earlier work this paper cites.
Dziugaite, G. K., and Roy, D. M · 2017
Earlier work this paper cites.
Learning neural networks with two nonlinear layers in polynomial time
Goel, S., and Klivans, A · 2017
Earlier work this paper cites.
Size-independent sample complexity of neural networks
Golowich, N., Rakhlin, A., and Shamir, O · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N · 2017
Cited alongside, same era.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N · 2017
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Sagun, L., Evci, U., Guney, V. U., Dauphin, Y., and Bottou, L · 2017
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Allen-Zhu, Z., Li, Y., and Liang, Y · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
A mean field view of the landscape of two-layers neural networks
Song, M., Montanari, A., and Nguyen, P · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Later among the works it cites.
Rademacher complexity for adversarially robust generalization
Yin, D., Ramchandran, K., and Bartlett, P · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2018
Later among the works it cites.
Arora, S., Du, S. S., Hu, W., Li, Z., and Wang, R · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Allen-Zhu, Z., Li, Y., and Song, Z · 2018
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
Arora, S., Ge, R., Neyshabur, B., and Zhang, Y · 2018
Cited alongside, same era.
Reconciling modern machine learning and the bias-variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2018
Cited alongside, same era.
A note on lazy training in supervised differentiable programming
Chizat, L., and Bach, F · 2018
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L., and Bach, F · 2018
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S. S., Lee, J. D., Li, H., Wang, L., and Zhai, X · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2018
Cited alongside, same era.
Learning one convolutional layer with overlapping patches
Goel, S., Klivans, A., and Meka, R · 2018
Cited alongside, same era.
Two models of double descent for weak features
Belkin, M., Hsu, D., and Xu, J · 2019
Closest in time.
A generalization theory of gradient descent for learning over-parameterized deep relu networks
Cao, Y., and Gu, Q · 2019
Closest in time.
Dou, X., and Liang, T · 2019
Closest in time.
Linearized two-layers neural networks in high dimension
Ghorbani, B., Mei, S., Misiakiewicz, T., and Montanari, A · 2019
Closest in time.
Li, M., Soltanolkotabi, M., and Oymak, S · 2019
Closest in time.
Size-free generalization bounds for convolutional neural networks
Long, P. M., and Sedghi, H · 2019
Closest in time.
Ma, C., Wu, L., et al · 2019
Closest in time.
Deterministic pac-bayesian generalization bounds for deep networks via generalizing noise-resilience
Nagarajan, V., and Kolter, J. Z · 2019
Closest in time.
Nitanda, A., and Suzuki, T · 2019
Closest in time.
Oymak, S., and Soltanolkotabi, M · 2019
Closest in time.
On the power and limitations of random features for understanding neural networks
Yehudai, G., and Shamir, O · 2019
Closest in time.
Training over-parameterized deep resnet is almost as easy as training a two-layer network
Zhang, H., Yu, D., Chen, W., and Liu, T.-Y · 2019
Closest in time.