Fetching the paper…
Reading the bibliography…
We study the first gradient descent step on the first-layer parameters $\boldsymbol{W}$ in a two-layer neural network: $f(\boldsymbol{x}) = \frac{1}{\sqrt{N}}\boldsymbol{a}^\top\sigma(\boldsymbol{W}^\top\boldsymbol{x})$, where $\boldsymbol{W}\in\mathbb{R}^{d\times N}, \boldsymbol{a}\in\mathbb{R}^{N}$ are randomly initialized, and the training objective is the empirical MSE loss: $\frac{1}{n}\sum_{i=1}^n (f(\boldsymbol{x}_i)-y_i)^2$.
On milman’s inequality and random subspaces which escape through a mesh in ℝ n \mathbb{R}^{n}
Yehoram Gordon · 1988
Earlier work this paper cites.
Matrix perturbation theory
Gilbert W Stewart · 1990
Earlier work this paper cites.
Bayesian learning for neural networks
Radford M Neal · 1995
Earlier work this paper cites.
No eigenvalues outside the support of the limiting spectral distribution of large-dimensional sample covariance matrices
Zhi-Dong Bai and Jack W Silverstein · 1998
Earlier work this paper cites.
On kernel-target alignment
Nello Cristianini, John Shawe-Taylor, Andre Elisseeff, and Jaz Kandola · 2001
Earlier work this paper cites.
Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices
Jinho Baik, Gérard Ben Arous, and Sandrine Péché · 2005
Earlier work this paper cites.
Spectra of large block matrices
Reza Rashidi Far, Tamer Oraby, Wlodzimierz Bryc, and Roland Speicher · 2006
Earlier work this paper cites.
Operator-valued semicircular elements: solving a quadratic matrix equation with positivity constraints
J William Helton, Reza Rashidi Far, and Roland Speicher · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Spectral analysis of large dimensional random matrices
Zhidong Bai and Jack W Silverstein · 2010
Earlier work this paper cites.
The spectrum of kernel random matrices
Noureddine El Karoui · 2010
Earlier work this paper cites.
The eigenvalues and eigenvectors of finite, low rank perturbations of large random matrices
Florent Benaych-Georges and Raj Rao Nadakuditi · 2011
Earlier work this paper cites.
The singular values and vectors of low rank perturbations of large rectangular random matrices
Florent Benaych-Georges and Raj Rao Nadakuditi · 2012
Earlier work this paper cites.
Concentration inequalities: A nonasymptotic theory of independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Earlier work this paper cites.
The spectrum of random inner-product kernel matrices
Xiuyuan Cheng and Amit Singer · 2013
Earlier work this paper cites.
The spectrum of random kernel matrices: universality results for rough and varying kernels
Yen Do and Van Vu · 2013
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik · 2014
Earlier work this paper cites.
A note on the hanson-wright inequality for random vectors with dependencies
Radoslaw Adamczak · 2015
Earlier work this paper cites.
Regularized linear regression: A precise analysis of the estimation error
Christos Thrampoulidis, Samet Oymak, and Babak Hassibi · 2015
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Earlier work this paper cites.
Free probability and random matrices
James A Mingo and Roland Speicher · 2017
Earlier work this paper cites.
Stochastic particle gradient descent for infinite ensembles
Atsushi Nitanda and Taiji Suzuki · 2017
Earlier work this paper cites.
Nonlinear random matrix theory for deep learning
Jeffrey Pennington and Pratik Worah · 2017
Earlier work this paper cites.
Limiting eigenvectors of outliers for spiked information-plus-noise type matrices
Mireille Capitaine · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
High-dimensional asymptotics of prediction: Ridge regression and classification
Edgar Dobriban and Stefan Wager · 2018
Earlier work this paper cites.
On the impact of predictor geometry on the performance on high-dimensional ridge-regularized generalized robust regression estimators
Noureddine El Karoui · 2018
Earlier work this paper cites.
Applications of realizations (aka linearizations) to free probability
J William Helton, Tobias Mai, and Roland Speicher · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
A random matrix approach to neural networks
Cosme Louart, Zhenyu Liao, and Romain Couillet · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Earlier work this paper cites.
Taiji Suzuki · 2018
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science
Roman Vershynin · 2018
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Earlier work this paper cites.
What can resnet learn efficiently, going beyond kernels?
Zeyuan Allen-Zhu and Yuanzhi Li · 2019
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Cited alongside, same era.
The spectral norm of random inner-product kernel matrices
Zhou Fan and Andrea Montanari · 2019
Cited alongside, same era.
Time matters in regularizing deep networks: Weight decay and data augmentation affect early learning dynamics, matter little near convergence
Aditya Sharad Golatkar, Alessandro Achille, and Stefano Soatto · 2019
Cited alongside, same era.
Limitations of lazy training of two-layers neural network
Taiji Suzuki and Shunta Akiyama · 2020
Later among the works it cites.
Nonparametric regression using deep neural networks with relu activation function
Johannes Schmidt-Hieber · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
On the optimal weighted ℓ 2 \ell_{2} regularization in overparameterized linear regression
Denny Wu and Ji Xu · 2020
Later among the works it cites.
Tensor programs iii: Neural matrix laws
Greg Yang · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Cited alongside, same era.
Deep neural networks learn non-smooth functions effectively
Masaaki Imaizumi and Kenji Fukumizu · 2019
Cited alongside, same era.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Cited alongside, same era.
The random matrix theory of the classical compact groups
Elizabeth S Meckes · 2019
Cited alongside, same era.
A note on the pennington-worah distribution
S Péché · 2019
Cited alongside, same era.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2019
Cited alongside, same era.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Cited alongside, same era.
Greg Yang and Edward J Hu · 2020
Later among the works it cites.
The staircase property: How hierarchical structure can guide deep learning
Emmanuel Abbe, Enric Boix Adsera, Matthew Brennan, Guy Bresler, and Dheeraj Nagaraj · 2021
Later among the works it cites.
Online stochastic gradient descent on non-convex losses from high-dimensional inference
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2021
Later among the works it cites.
Model, sample, and epoch-wise descents: exact solution of gradient flow in the random feature model
Antoine Bodin and Nicolas Macris · 2021
Later among the works it cites.
Deep learning: a statistical viewpoint
Peter L Bartlett, Andrea Montanari, and Alexander Rakhlin · 2021
Later among the works it cites.
Eigenvalue distribution of some nonlinear models of random matrices
Lucas Benigni and Sandrine Péché · 2021
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
Later among the works it cites.
When does gradient descent with logistic loss find interpolating two-layer networks?
Niladri S Chatterji, Philip M Long, and Peter L Bartlett · 2021
Later among the works it cites.
How rotational invariance of common kernels prevents generalization in high dimensions
Konstantin Donhauser, Mingqi Wu, and Fanny Yang · 2021
Later among the works it cites.
The gaussian equivalence of generative models for learning with shallow neural networks
Sebastian Goldt, Bruno Loureiro, Galen Reeves, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2021
Later among the works it cites.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
Later among the works it cites.
Training integrable parameterizations of deep neural networks in the infinite-width limit
Karl Hajjar, Lénaïc Chizat, and Christophe Giraud · 2021
Later among the works it cites.
Local signal adaptivity: Provable feature learning in neural networks beyond kernels
Stefani Karp, Ezra Winston, Yuanzhi Li, and Aarti Singh · 2021
Later among the works it cites.
Learning curves of generic features maps for realistic datasets with a teacher-student model
Bruno Loureiro, Cedric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mezard, and Lenka Zdeborová · 2021
Later among the works it cites.
Quantifying the benefit of using differentiable learning over tangent kernels
Eran Malach, Pritish Kamath, Emmanuel Abbe, and Nathan Srebro · 2021
Later among the works it cites.
Generalization error of random feature and kernel methods: hypercontractivity and kernel matrix concentration
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
Later among the works it cites.
Analysis of feature learning in weight-tied autoencoders via the mean field lens
Phan-Minh Nguyen · 2021
Later among the works it cites.
What can linearized neural networks actually say about generalization?
Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli, and Pascal Frossard · 2021
Later among the works it cites.
Implicit bias of sgd for diagonal linear networks: a provable benefit of stochasticity
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Later among the works it cites.
Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed
Maria Refinetti, Sebastian Goldt, Florent Krzakala, and Lenka Zdeborová · 2021
Later among the works it cites.
Covariate shift in high-dimensional random feature regression
Nilesh Tripuraneni, Ben Adlam, and Jeffrey Pennington · 2021
Later among the works it cites.
Zhichao Wang and Yizhe Zhu · 2021
Later among the works it cites.
Emmanuel Abbe, Enric Boix-Adsera, and Theodor Misiakiewicz · 2022
Closest in time.
Largest eigenvalues of the conjugate kernel of single-layered neural networks
Lucas Benigni and Sandrine Péché · 2022
Closest in time.
Mean-field langevin dynamics: Exponential convergence and annealing
Lénaïc Chizat · 2022
Closest in time.
Random feature amplification: Feature learning and generalization in neural networks
Spencer Frei, Niladri S Chatterji, and Peter L Bartlett · 2022
Closest in time.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari · 2022
Closest in time.
Universality of empirical risk minimization
Andrea Montanari and Basil Saeed · 2022
Closest in time.
Convex analysis of the mean field langevin dynamics
Atsushi Nitanda, Denny Wu, and Taiji Suzuki · 2022
Closest in time.
Phase diagram of stochastic gradient descent in high-dimensional two-layer neural networks
Rodrigo Veiga, Ludovic Stephan, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2022
Closest in time.
Learning Theory from First Principles
Francis Bach · 2023
Closest in time.