Fetching the paper…
Reading the bibliography…
One of the distinguishing characteristics of modern deep learning systems is that they typically employ neural network architectures that utilize enormous numbers of parameters, often in the millions and sometimes even in the billions.
Distribution of eigenvalues for some sets of random matrices
V. A. Marčenko and L. A. Pastur · 1967
Earlier work this paper cites.
Generalized cross-validation as a method for choosing a good ridge parameter
G. H. Golub, M. Heath, and G. Wahba · 1979
Earlier work this paper cites.
On the empirical distribution of eigenvalues of a class of large dimensional random matrices
J. W. Silverstein and Z. Bai · 1995
Earlier work this paper cites.
On the empirical distribution of eigenvalues of a class of large dimensional random matrices, 1995
J. W. Silverstein and Z. D. Bai · 1995
Earlier work this paper cites.
Spectral moments of correlated wishart matrices
Z. Burda, J. Jurkiewicz, and B. Wacław · 2005
Earlier work this paper cites.
Spectral analysis of large dimensional random matrices
Z. Lixin · 2006
Earlier work this paper cites.
Large sample covariance matrices without independence structures in columns
Z. Bai and W. Zhou · 2008
Earlier work this paper cites.
Concentration of measure and spectra of random matrices: Applications to correlation matrices, elliptical distributions and beyond
N. El Karoui et al · 2009
Earlier work this paper cites.
No eigenvalues outside the support of the limiting empirical spectral distribution of a separable covariance matrix
D. Paul and J. W. Silverstein · 2009
Earlier work this paper cites.
The spectrum of kernel random matrices
N. El Karoui · 2010
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Topics in random matrix theory
T. Tao · 2012
Earlier work this paper cites.
Random matrix theory in statistics: A review
D. Paul and A. Aue · 2014
Earlier work this paper cites.
The universality principle for spectral distributions of sample covariance matrices
P. Yaskov · 2014
Earlier work this paper cites.
Limiting spectral distribution of large sample covariance matrices associated with a class of stationary processes
M. Banna and F. Merlevede · 2015
Cited alongside, same era.
On the limiting spectral distribution for a large class of symmetric random matrices with correlated entries
M. Banna, F. Merlevède, and M. Peligrad · 2015
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al · 2016
Cited alongside, same era.
A dynamical approach to random matrix theory
L. Erdos and H.-T. Yau · 2017
Cited alongside, same era.
Geometry of neural network loss surfaces via random matrix theory
J. Pennington and Y. Bahri · 2017
Cited alongside, same era.
The spectrum of the fisher information matrix of a single-hidden-layer neural network
J. Pennington and P. Worah · 2018
Later among the works it cites.
G. M. Rotskoff and E. Vanden-Eijnden · 2018
Later among the works it cites.
A note on lazy training in supervised differentiable programming
L. Chizat and F. Bach · 2019
Closest in time.
The matrix dyson equation and its applications for random matrices
L. Erdos · 2019
Closest in time.
Surprises in high-dimensional ridgeless least squares interpolation
T. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nonlinear random matrix theory for deep learning
J. Pennington and P. Worah · 2017
Cited alongside, same era.
Outrageously large neural language models using sparsely gated mixtures of experts
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
L. Chizat and F. Bach · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Cited alongside, same era.
An analytic theory of generalization dynamics and transfer learning in deep linear networks
A. K. Lampinen and S. Ganguli · 2018
Cited alongside, same era.
The dynamics of learning: A random matrix approach
Z. Liao and R. Couillet · 2018
Cited alongside, same era.
On the Spectrum of Random Features Maps of High Dimensional Data
Z. Liao and R. Couillet · 2018
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
J. Lee, L. Xiao, S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington · 2019
Closest in time.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
S. Mei, T. Misiakiewicz, and A. Montanari · 2019
Closest in time.
The generalization error of random features regression: Precise asymptotics and the double descent curve
S. Mei and A. Montanari · 2019
Closest in time.
Global convergence of neuron birth-death dynamics
G. Rotskoff, S. Jelassi, J. Bruna, and E. Vanden-Eijnden · 2019
Closest in time.
The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization
B. Adlam and J. Pennington · 2020
Closest in time.
Understanding double descent requires a fine-grained bias-variance decomposition
B. Adlam and J. Pennington · 2020
Closest in time.
High-dimensional dynamics of generalization error in neural networks
M. S. Advani, A. M. Saxe, and H. Sompolinsky · 2020
Closest in time.
Double trouble in double descent: Bias and variance (s) in the lazy regime
S. d?Ascoli, M. Refinetti, G. Biroli, and F. Krzakala · 2020
Closest in time.
Kernel alignment risk estimator: Risk prediction from training data
A. Jacot, B. Simsek, F. Spadaro, C. Hongler, and F. Gabriel · 2020
Closest in time.