Fetching the paper…
Reading the bibliography…
Modern deep learning models employ considerably more parameters than required to fit the training data.
Two models of double descent for weak features
Belkin, M., Hsu, D., and Xu, J · 1903
Earlier work this paper cites.
On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels
Liang, T., Rakhlin, A., and Zhai, X · 1908
Earlier work this paper cites.
Generalized cross-validation as a method for choosing a good ridge parameter
Golub, G. H., Heath, M., and Wahba, G · 1979
Earlier work this paper cites.
Priors for infinite networks
Neal, R. M · 1996
Earlier work this paper cites.
Spectra of large block matrices
Far, R. R., Oraby, T., Bryc, W., and Speicher, R · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B · 2008
Earlier work this paper cites.
The spectrum of kernel random matrices
El Karoui, N. et al · 2010
Earlier work this paper cites.
An introduction to matrix concentration inequalities
Tropp, J. A · 2015
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Earlier work this paper cites.
High-dimensional dynamics of generalization error in neural networks
Advani, M. S. and Saxe, A. M · 2017
Earlier work this paper cites.
Free probability and random matrices , volume 35
Mingo, J. A. and Speicher, R · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N · 2017
Earlier work this paper cites.
Nonlinear random matrix theory for deep learning
Pennington, J. and Worah, P · 2017
Earlier work this paper cites.
Gaussian process behaviour in wide deep neural networks
de G. Matthews, A. G., Hron, J., Rowland, M., Turner, R. E., and Ghahramani, Z · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2018
Cited alongside, same era.
Applications of realizations (aka linearizations) to free probability
Helton, J. W., Mai, T., and Speicher, R · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Deep neural networks as gaussian processes
Lee, J., Bahri, Y., Novak, R., Schoenholz, S. S., Pennington, J., and Sohl-Dickstein, J · 2018
Cited alongside, same era.
A random matrix approach to neural networks
Louart, C., Liao, Z., Couillet, R., et al · 2018
The matrix dyson equation and its applications for random matrices
Erdos, L · 2019
Later among the works it cites.
Linearized two-layers neural networks in high dimension
Ghorbani, B., Mei, S., Misiakiewicz, T., and Montanari, A · 2019
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
Hastie, T., Montanari, A., Rosset, S., and Tibshirani, R. J · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The spectrum of the fisher information matrix of a single-hidden-layer neural network
Pennington, J. and Worah, P · 2018
Cited alongside, same era.
A random matrix perspective on mixtures of nonlinearities for deep learning
Adlam, B., Levinson, J., and Pennington, J · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R. R., and Wang, R · 2019
Cited alongside, same era.
Eigenvalue distribution of nonlinear models of random matrices
Benigni, L. and Péché, S · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S., Lee, J., Li, H., Wang, L., and Zhai, X · 2019
Cited alongside, same era.
Li, Z., Wang, R., Yu, D., Du, S. S., Hu, W., Salakhutdinov, R., and Arora, S · 2019
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Mei, S. and Montanari, A · 2019
Later among the works it cites.
Mitra, P. P · 2019
Later among the works it cites.
Uniform convergence may be unable to explain generalization in deep learning
Nagarajan, V. and Kolter, J. Z · 2019
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 2019
Later among the works it cites.
A note on the pennington-worah distribution
Péché, S. et al · 2019
Later among the works it cites.
Disentangling trainability and generalization in deep learning
Xiao, L., Pennington, J., and Schoenholz, S. S · 2019
Later among the works it cites.
Scaling description of generalization with number of parameters in deep learning
Geiger, M., Jacot, A., Spigler, S., Gabriel, F., Sagun, L., d’Ascoli, S., Biroli, G., Hongler, C., and Wyart, M · 2020
Closest in time.