Fetching the paper…
Reading the bibliography…
We study the eigenvalue distributions of the Conjugate Kernel and Neural Tangent Kernel associated to multi-layer feedforward neural networks.
Distribution of eigenvalues for some sets of random matrices
Vladimir Alexandrovich Marchenko and Leonid Andreevich Pastur · 1967
Earlier work this paper cites.
Matrix Theory and Applications
C.R. Johnson · 1990
Earlier work this paper cites.
Bayesian learning for neural networks
Radford M Neal · 1995
Earlier work this paper cites.
Computing with infinite networks
Christopher KI Williams · 1997
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence K Saul · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
On a surprising relation between the marchenko-pastur law, rectangular and square free convolutions
Florent Benaych-Georges · 2010
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
Concentration Inequalities: A Nonasymptotic Theory of Independence
S. Boucheron, G. Lugosi, and P. Massart · 2013
Earlier work this paper cites.
Hanson-Wright inequality and sub-gaussian concentration
Mark Rudelson and Roman Vershynin · 2013
Earlier work this paper cites.
A note on the Hanson-Wright inequality for random vectors with dependencies
Radoslaw Adamczak · 2015
Earlier work this paper cites.
Concentration inequalities for non-Lipschitz functions with bounded derivatives of higher order
Radosław Adamczak and Paweł Wolff · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Random weighted projections, random quadratic forms and random eigenvectors
Van Vu and Ke Wang · 2015
Earlier work this paper cites.
Kernel spectral clustering of large dimensional data
Romain Couillet and Florent Benaych-Georges · 2016
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Ridge regression and asymptotic minimax estimation over spheres of growing dimension
Lee H Dicker · 2016
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Earlier work this paper cites.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani and Andrew M Saxe · 2017
Earlier work this paper cites.
Geometry of neural network loss surfaces via random matrix theory
Jeffrey Pennington and Yasaman Bahri · 2017
Cited alongside, same era.
Nonlinear random matrix theory for deep learning
Jeffrey Pennington and Pratik Worah · 2017
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Cited alongside, same era.
Deep information propagation
Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
High-dimensional asymptotics of prediction: Ridge regression and classification
Edgar Dobriban and Stefan Wager · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Asymptotics of wide networks from Feynman diagrams
Ethan Dyer and Guy Gur-Ari · 2019
Later among the works it cites.
Jamming transition as a paradigm to understand the loss landscape of deep neural networks
Mario Geiger, Stefano Spigler, Stéphane d’Ascoli, Levent Sagun, Marco Baity-Jesi, Giulio Biroli, and Matthieu Wyart · 2019
Later among the works it cites.
Limitations of lazy training of two-layers neural network
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Later among the works it cites.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep neural networks as Gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
The dynamics of learning: A random matrix approach
Zhenyu Liao and Romain Couillet · 2018
Cited alongside, same era.
On the spectrum of random features maps of high dimensional data
Zhenyu Liao and Romain Couillet · 2018
Cited alongside, same era.
A random matrix approach to neural networks
Cosme Louart, Zhenyu Liao, and Romain Couillet · 2018
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
Alexander G de G Matthews, Jiri Hron, Mark Rowland, Richard E Turner, and Zoubin Ghahramani · 2018
Cited alongside, same era.
The spectrum of the Fisher information matrix of a single-hidden-layer neural network
Jeffrey Pennington and Pratik Worah · 2018
Cited alongside, same era.
Later among the works it cites.
Dynamics of deep neural networks and neural tangent hierarchy
Jiaoyang Huang and Horng-Tzer Yau · 2019
Later among the works it cites.
The asymptotic spectrum of the hessian of dnn throughout training
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2019
Later among the works it cites.
Universal statistics of Fisher information in deep neural networks: Mean field approach
Ryo Karakida, Shotaro Akaho, and Shun-ichi Amari · 2019
Later among the works it cites.
Restricted isometry property under high correlations
Shiva Prasad Kasiviswanathan and Mark Rudelson · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Later among the works it cites.
On the risk of minimum-norm interpolants and restricted lower isometry of kernels
Tengyuan Liang, Alexander Rakhlin, and Xiyu Zhai · 2019
Later among the works it cites.
On inner-product kernels of high dimensional data
Zhenyu Liao and Romain Couillet · 2019
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Later among the works it cites.
A note on the pennington-worah distribution
S Péché · 2019
Later among the works it cites.
Disentangling trainability and generalization in deep learning
Lechao Xiao, Jeffrey Pennington, and Samuel S Schoenholz · 2019
Later among the works it cites.
Greg Yang · 2019
Later among the works it cites.
A fine-grained spectral perspective on neural networks
Greg Yang and Hadi Salman · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Later among the works it cites.
The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization
Ben Adlam and Jeffrey Pennington · 2020
Closest in time.
Double trouble in double descent: Bias and variance(s) in the lazy regime
Stéphane d’Ascoli, Maria Refinetti, Giulio Biroli, and Florent Krzakala · 2020
Closest in time.