Fetching the paper…
Reading the bibliography…
In this paper, we investigate a two-layer fully connected neural network of the form $f(X)=\frac{1}{\sqrt{d_1}}\boldsymbol{a}^\top \sigma\left(WX\right)$, where $X\in\mathbb{R}^{d_0\times n}$ is a deterministic data matrix, $W\in\mathbb{R}^{d_1\times d_0}$ and $\boldsymbol{a}\in\mathbb{R}^{d_1}$ are random Gaussian weights, and $\sigma$ is a nonlinear activation function.
An inverse matrix adjustment arising in discriminant analysis
Maurice S Bartlett · 1951
Earlier work this paper cites.
A bound on tail probabilities for quadratic forms in independent random variables
David Lee Hanson and Farroll Tim Wright · 1971
Earlier work this paper cites.
The smallest eigenvalue of a large dimensional wishart matrix
Jack W Silverstein · 1985
Earlier work this paper cites.
Multiplication of certain non-commuting random variables
Dan Voiculescu · 1987
Earlier work this paper cites.
Convergence to the semicircle law
Zhidong Bai and Y. Q. Yin · 1988
Earlier work this paper cites.
Matrix Theory and Applications
C.R. Johnson · 1990
Earlier work this paper cites.
Bayesian learning for neural networks
Radford M Neal · 1995
Earlier work this paper cites.
Computing with infinite networks
Christopher KI Williams · 1997
Earlier work this paper cites.
The limiting distributions of eigenvalues of sample correlation matrices
Tiefeng Jiang · 2004
Earlier work this paper cites.
Lectures on the combinatorics of free probability
Alexandru Nica and Roland Speicher · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence K Saul · 2009
Earlier work this paper cites.
An introduction to random matrices
Greg W Anderson, Alice Guionnet, and Ofer Zeitouni · 2010
Earlier work this paper cites.
Spectral analysis of large dimensional random matrices
Zhidong Bai and Jack W Silverstein · 2010
Earlier work this paper cites.
The limiting spectral distribution of the product of the Wigner matrix and a nonnegative definite matrix
ZD Bai and LX Zhang · 2010
Earlier work this paper cites.
Partial transposition of random states and non-centered semicircular distributions
Guillaume Aubrun · 2012
Earlier work this paper cites.
Strong convergence of esd for the generalized sample covariance matrices when p / n → 0 p/n\to 0
Zhigang Bao · 2012
Earlier work this paper cites.
Convergence of the largest eigenvalue of normalized sample covariance matrices when p p and n n both tend to infinity with their ratio converging to zero
Binbin Chen and Guangming Pan · 2012
Earlier work this paper cites.
User-friendly tail bounds for sums of random matrices
Joel A Tropp · 2012
Earlier work this paper cites.
Sharp analysis of low-rank kernel matrix approximations
Francis Bach · 2013
Earlier work this paper cites.
Concentration inequalities: A nonasymptotic theory of independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Earlier work this paper cites.
Hanson-wright inequality and sub-gaussian concentration
Mark Rudelson and Roman Vershynin · 2013
Earlier work this paper cites.
Limiting spectral distribution of normalized sample covariance matrices with p / n → 0 p/n\to 0
Junshan Xie · 2013
Earlier work this paper cites.
Limiting spectral distribution of renormalized separable sample covariance matrices when p / n → 0 p/n\to 0
Lili Wang and Debashis Paul · 2014
Earlier work this paper cites.
A note on the Hanson-Wright inequality for random vectors with dependencies
Radoslaw Adamczak · 2015
Earlier work this paper cites.
CLT for linear spectral statistics of normalized sample covariance matrices with the dimension much larger than the sample size
Binbin Chen and Guangming Pan · 2015
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Testing the sphericity of a covariance matrix when the dimension is much larger than the sample size
Zeng Li and Jianfeng Yao · 2016
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Earlier work this paper cites.
Random fourier features for kernel ridge regression: Approximation bounds and statistical guarantees
Haim Avron, Michael Kapralov, Cameron Musco, Christopher Musco, Ameya Velingker, and Amir Zandieh · 2017
Cited alongside, same era.
Alice and Bob meet Banach
Guillaume Aubrun and Stanisław J Szarek · 2017
Cited alongside, same era.
On the equivalence between kernel quadrature rules and random feature expansions
Francis Bach · 2017
Cited alongside, same era.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2017
Cited alongside, same era.
Nonlinear random matrix theory for deep learning
Jeffrey Pennington and Pratik Worah · 2017
Cited alongside, same era.
Generalization properties of learning with random features
Alessandro Rudi and Lorenzo Rosasco · 2017
Generalisation error in learning with random features and the hidden manifold model
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2020
Later among the works it cites.
The surprising simplicity of the early-time learning dynamics of neural networks
Wei Hu, Lechao Xiao, Ben Adlam, and Jeffrey Pennington · 2020
Later among the works it cites.
Implicit regularization of random feature models
Arthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler, and Franck Gabriel · 2020
Later among the works it cites.
A random matrix analysis of random fourier features: beyond the gaussian kernel, a precise phase transition, and the corresponding double descent
Zhenyu Liao, Romain Couillet, and Michael W. Mahoney · 2020
Later among the works it cites.
Just interpolate: Kernel “ridgeless” regression can generalize
Tengyuan Liang and Alexander Rakhlin · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep information propagation
Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
The PPT square conjecture holds generically for some classes of independent states
Benoît Collins, Zhi Yin, and Ping Zhong · 2018
Cited alongside, same era.
Neural tangent kernel: convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Deep neural networks as Gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
On the spectrum of random features maps of high dimensional data
Zhenyu Liao and Romain Couillet · 2018
Cited alongside, same era.
A random matrix approach to neural networks
Cosme Louart, Zhenyu Liao, and Romain Couillet · 2018
Cited alongside, same era.
On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels
Tengyuan Liang, Alexander Rakhlin, and Xiyu Zhai · 2020
Later among the works it cites.
Global convergence of deep networks with one wide layer followed by pyramidal topology
Quynh Nguyen and Marco Mondelli · 2020
Later among the works it cites.
Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
Deep learning: a statistical viewpoint
Peter L Bartlett, Andrea Montanari, and Alexander Rakhlin · 2021
Closest in time.
Eigenvalue distribution of some nonlinear models of random matrices
Lucas Benigni and Sandrine Péché · 2021
Closest in time.
Spiked singular values and vectors under extreme aspect ratios
Michael J Feldman · 2021
Closest in time.
Large-dimensional random matrix theory and its applications in deep learning and wireless communications
Jungang Ge, Ying-Chang Liang, Zhidong Bai, and Guangming Pan · 2021
Closest in time.
The spectrum of Fisher information of deep networks achieving dynamical isometry
Tomohiro Hayase and Ryo Karakida · 2021
Closest in time.
What causes the test error? going beyond bias-variance via anova
Licong Lin and Edgar Dobriban · 2021
Closest in time.
Learning curves of generic features maps for realistic datasets with a teacher-student model
Bruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2021
Closest in time.
Kernel regression in high dimensions: Refined analysis beyond double descent
Fanghui Liu, Zhenyu Liao, and Johan Suykens · 2021
Closest in time.
Generalization error of random feature and kernel methods: hypercontractivity and kernel matrix concentration
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
Closest in time.
On the proof of global convergence of gradient descent for deep relu networks with linear widths
Quynh Nguyen · 2021
Closest in time.
Tight bounds on the smallest eigenvalue of the neural tangent kernel for deep relu networks
Quynh Nguyen, Marco Mondelli, and Guido F Montufar · 2021
Closest in time.
Analysis of one-hidden-layer neural networks via the resolvent method
Vanessa Piccolo and Dominik Schröder · 2021
Closest in time.
Jiaxin Qiu, Zeng Li, and Jianfeng Yao · 2021
Closest in time.
Exact gap between generalization error and uniform convergence in random feature models
Zitong Yang, Yu Bai, and Song Mei · 2021
Closest in time.
A random matrix perspective on mixtures of nonlinearities in high dimensions
Ben Adlam, Jake A Levinson, and Jeffrey Pennington · 2022
Closest in time.
Learning rates as a function of batch size: A random matrix theory approach to neural network training
Diego Granziol, Stefan Zohren, and Stephen Roberts · 2022
Closest in time.
Universality laws for high-dimensional learning with random features
Hong Hu and Yue M Lu · 2022
Closest in time.
The interpolation phase transition in neural networks: Memorization and generalization under lazy training
Andrea Montanari and Yiqiao Zhong · 2022
Closest in time.
Overparameterized random feature regression with nearly orthogonal data
Zhichao Wang and Yizhe Zhu · 2022
Closest in time.
Testing Kronecker product covariance matrices for high-dimensional matrix-variate data
Long Yu, Jiahui Xie, and Wang Zhou · 2022
Closest in time.
Asymptotic freeness of layerwise jacobians caused by invariance of multilayer perceptron: The haar orthogonal case
Benoit Collins and Tomohiro Hayase · 2023
Closest in time.