Fetching the paper…
Reading the bibliography…
For a large class of feature maps we provide a tight asymptotic characterisation of the test error associated with learning the readout layer, in the high-dimensional limit where the input dimension, hidden layer widths, and number of training samples are proportionally large.
Distribution of eigenvalues for some sets of random matrices
V A Marčenko and L A Pastur · 1967
Earlier work this paper cites.
Strong convergence of the empirical distribution of eigenvalues of large dimensional random matrices
J.W. Silverstein · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Patrick Haffner, Yoshua Bengio, Yann LeCun, Yann LeCun, and Léon Bottou · 1998
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Large sample covariance matrices without independence structures in columns
Zhidong Bai and Wang Zhou · 2008
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
Spectral convergence for a general class of random matrices
Francisco Rubio and Xavier Mestre · 2011
Earlier work this paper cites.
Normal approximations with Malliavin calculus: from Stein’s method to universality
Ivan Nourdin and Giovanni Peccati · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
A note on the hanson-wright inequality for random vectors with dependencies
Radosław Adamczak · 2015
Earlier work this paper cites.
A note on the Hanson-Wright inequality for random vectors with dependencies
Radosł aw Adamczak · 2015
Earlier work this paper cites.
Anisotropic local laws for random matrices
Antti Knowles and Jun Yin · 2017
Earlier work this paper cites.
On the spectrum of random features maps of high dimensional data
Zhenyu Liao and Romain Couillet · 2018
Earlier work this paper cites.
High-dimensional asymptotics of prediction: ridge regression and classification
Edgar Dobriban and Stefan Wager · 2018
Earlier work this paper cites.
A random matrix approach to neural networks
Cosme Louart, Zhenyu Liao, and Romain Couillet · 2018
Earlier work this paper cites.
A random matrix approach to neural networks
Cosme Louart, Zhenyu Liao, and Romain Couillet · 2018
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Siyuan Ma, Soumik Mandal, Mikhail Belkin, and Daniel Hsu · 2019
Earlier work this paper cites.
Nonlinear random matrix theory for deep learning
Jeffrey Pennington and Pratik Worah · 2019
Earlier work this paper cites.
Limitations of lazy training of two-layers neural network
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Earlier work this paper cites.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Earlier work this paper cites.
Benign overfitting in linear regression
Peter L. Bartlett, Philip M. Long, Gábor Lugosi, and Alexander Tsigler · 2020
Earlier work this paper cites.
Modeling the Influence of Data Structure on Learning in Neural Networks: The Hidden Manifold Model
Sebastian Goldt, Marc Mézard, Florent Krzakala, and Lenka Zdeborová · 2020
Cited alongside, same era.
A precise performance analysis of learning with random features
Oussama Dhifallah and Yue M Lu · 2020
Cited alongside, same era.
Spectra of the conjugate kernel and neural tangent kernel for linear-width neural networks
Zhou Fan and Zhichao Wang · 2020
Cited alongside, same era.
The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization
Ben Adlam and Jeffrey Pennington · 2020
Cited alongside, same era.
Implicit regularization of random feature models
Arthur Jacot, Berfin Simsek, Francesco Spadaro, Clement Hongler, and Franck Gabriel · 2020
Cited alongside, same era.
Learning curves of generic features maps for realistic datasets with a teacher-student model
Bruno Loureiro, Cedric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2022
Later among the works it cites.
Dimension free ridge regression
Chen Cheng and Andrea Montanari · 2022
Later among the works it cites.
Clément Chouard · 2022
Later among the works it cites.
Rank-uniform local law for Wigner matrices
Giorgio Cipolloni, László Erdős, and Dominik Schröder · 2022
Later among the works it cites.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
Jimmy Ba, Murat A Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, and Greg Yang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the optimal weighted ℓ _ 2 \ell\_2 regularization in overparameterized linear regression
Denny Wu and Ji Xu · 2020
Cited alongside, same era.
The Gaussian equivalence of generative models for learning with shallow neural networks
Sebastian Goldt, Bruno Loureiro, Galen Reeves, Florent Krzakala, Marc Mezard, and Lenka Zdeborová · 2021
Cited alongside, same era.
Generalisation error in learning with random features and the hidden manifold model
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2021
Cited alongside, same era.
Implicit self-regularization in deep neural networks: Evidence from random matrix theory and implications for learning
Charles H. Martin and Michael W. Mahoney · 2021
Cited alongside, same era.
Eigenvalue distribution of some nonlinear models of random matrices
Lucas Benigni and Sandrine Péché · 2021
Cited alongside, same era.
When do neural networks outperform kernel methods?
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
Cited alongside, same era.
Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed
Maria Refinetti, Sebastian Goldt, Florent Krzakala, and Lenka Zdeborová · 2021
Cited alongside, same era.
Deterministic equivalent and error universality of deep random features learning
Dominik Schröder, Hugo Cui, Daniil Dmitriev, and Bruno Loureiro · 2023
Later among the works it cites.
Precise asymptotic analysis of deep random feature models
David Bosch, Ashkan Panahi, and Babak Hassibi · 2023
Later among the works it cites.
Learning curves for deep structured gaussian feature models
Jacob Zavatone-Veth and Cengiz Pehlevan · 2023
Later among the works it cites.
A rainbow in deep network black boxes
Florentin Guth, Brice Ménard, Gaspar Rochette, and Stéphane Mallat · 2023
Later among the works it cites.
Bayes-optimal learning of deep random networks of extensive-width
Hugo Cui, Florent Krzakala, and Lenka Zdeborova · 2023
Later among the works it cites.
Random features model with general convex regularization: A fine grained analysis with precise asymptotic learning curves
David Bosch, Ashkan Panahi, Ayca Ozcelikkale, and Devdatt Dubhashi · 2023
Later among the works it cites.
Matrix dyson equation for correlated linearizations and test error of random features regression
Hugo Latourelle-Vigeant and Elliot Paquette · 2023
Later among the works it cites.
Deterministic equivalent of the conjugate kernel matrix associated to artificial neural networks
Clément Chouard · 2023
Later among the works it cites.
Learning two-layer neural networks, one (giant) step at a time
Yatin Dandi, Florent Krzakala, Bruno Loureiro, Luca Pesce, and Ludovic Stephan · 2023
Later among the works it cites.
A theory of non-linear feature learning with one gradient step in two-layer neural networks
Behrad Moniri, Donghwan Lee, Hamed Hassani, and Edgar Dobriban · 2023
Later among the works it cites.
Are Gaussian data all you need? The extents and limits of universality in high-dimensional generalized linear estimation
Luca Pesce, Florent Krzakala, Bruno Loureiro, and Ludovic Stephan · 2023
Later among the works it cites.
Gaussian universality of perceptrons with random labels
Federica Gerace, Florent Krzakala, Bruno Loureiro, Ludovic Stephan, and Lenka Zdeborová · 2024
Closest in time.
High-dimensional analysis of double descent for linear regression with random projections
Francis Bach · 2024
Closest in time.
Asymptotics of feature learning in two-layer networks after one gradient-step
Hugo Cui, Luca Pesce, Yatin Dandi, Florent Krzakala, Yue M Lu, Lenka Zdeborová, and Bruno Loureiro · 2024
Closest in time.
Gaussian universality of perceptrons with random labels
Federica Gerace, Florent Krzakala, Bruno Loureiro, Ludovic Stephan, and Lenka Zdeborová · 2024
Closest in time.