Fetching the paper…
Reading the bibliography…
Feature learning is thought to be one of the fundamental reasons for the success of deep neural networks.
Eine neue herleitung des exponentialgesetzes in der wahrscheinlichkeitsrechnung
Jarl Waldemar Lindeberg · 1922
Earlier work this paper cites.
Characteristic vectors of bordered matrices with infinite dimensions
Eugene P Wigner · 1955
Earlier work this paper cites.
Sur la loi limite de l’espacement des valeurs propres d’une matrice alé atoire
Michel Gaudin · 1961
Earlier work this paper cites.
A Brownian-motion model for the eigenvalues of a random matrix
Freeman J Dyson · 1962
Earlier work this paper cites.
On determining training sample size of linear classifier
Šarūnas Raudys · 1967
Earlier work this paper cites.
Handbook of mathematical functions with formulas, graphs, and mathematical tables
Milton Abramowitz and Irene A Stegun · 1968
Earlier work this paper cites.
Representation of statistics of discriminant analysis and asymptotic expansion when space dimensions are comparable with sample size
AD Deev · 1970
Earlier work this paper cites.
On the amount of a priori information in designing the classification algorithm
Šarūnas Raudys · 1972
Earlier work this paper cites.
Perturbation bounds in connection with singular value decomposition
Per-Ake Wedin · 1972
Earlier work this paper cites.
Statistical theory of learning a rule
Géza Györgyi and Naftali Tishby · 1990
Earlier work this paper cites.
Matrix perturbation theory
Gilbert W Stewart and Ji-guang Sun · 1990
Earlier work this paper cites.
Statistical mechanics of learning: Generalization
Manfred Opper · 1995
Earlier work this paper cites.
Statistical mechanics of generalization
Manfred Opper and Wolfgang Kinzel · 1996
Earlier work this paper cites.
Statistical mechanics of learning
Andreas Engel and Christian Van den Broeck · 2001
Earlier work this paper cites.
Random matrices
Madan Lal Mehta · 2004
Earlier work this paper cites.
Results in statistical discriminant analysis: A review of the former Soviet Union literature
Šarūnas Raudys and Dean M Young · 2004
Earlier work this paper cites.
Random matrix theory and wireless communications
Antonio M Tulino and Sergio Verdú · 2004
Earlier work this paper cites.
On the empirical distribution of eigenvalues of large dimensional information-plus-noise-type matrices
R Brent Dozier and Jack W Silverstein · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Multiparametric Statistics
Vadim Ivanovich Serdobolskii · 2007
Earlier work this paper cites.
Spectral Analysis of Large Dimensional Random Matrices
Zhidong Bai and Jack W Silverstein · 2010
Earlier work this paper cites.
The spectrum of kernel random matrices
Noureddine El Karoui · 2010
Earlier work this paper cites.
Bulk universality for Wigner matrices
László Erdös, Sandrine Péché, José A Ramírez, Benjamin Schlein, and Horng-Tzer Yau · 2010
Earlier work this paper cites.
Random Matrix Methods for Wireless Communications
Romain Couillet and Merouane Debbah · 2011
Earlier work this paper cites.
Random matrices: universality of local eigenvalue statistics
Terence Tao and Van Vu · 2011
Earlier work this paper cites.
Bulk universality for generalized Wigner matrices
László Erdös, Horng-Tzer Yau, and Jun Yin · 2012
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
The spectrum of random inner-product kernel matrices
Xiuyuan Cheng and Amit Singer · 2013
Earlier work this paper cites.
Weak convergence and empirical processes: with applications to statistics
Aad van der Vaart and Jon Wellner · 2013
Earlier work this paper cites.
Exact separation phenomenon for the eigenvalues of large information-plus-noise type matrices, and an application to spiked models
Mireille Capitaine · 2014
Earlier work this paper cites.
Analysis of boolean functions
Ryan O’Donnell · 2014
Earlier work this paper cites.
Random matrix theory in statistics: A review
Debashis Paul and Alexander Aue · 2014
Earlier work this paper cites.
On the limiting spectral distribution for a large class of symmetric random matrices with correlated entries
Marwa Banna, Florence Merlevède, and Magda Peligrad · 2015
Earlier work this paper cites.
Large Sample Covariance Matrices and High-Dimensional Data Analysis
Jianfeng Yao, Zhidong Bai, and Shurong Zheng · 2015
Earlier work this paper cites.
Adversarial feature learning
Jeff Donahue, Philipp Krähenbühl, and Trevor Darrell · 2016
Earlier work this paper cites.
Nonlinear random matrix theory for deep learning
Jeffrey Pennington and Pratik Worah · 2017
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Cited alongside, same era.
High-dimensional asymptotics of prediction: Ridge regression and classification
Edgar Dobriban and Stefan Wager · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
A random matrix approach to neural networks
Cosme Louart, Zhenyu Liao, and Romain Couillet · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Feature learning in infinite-width neural networks
Greg Yang and Edward J Hu · 2021
Later among the works it cites.
Concentration inequalities for statistical inference
Huiming Zhang and Songxi Chen · 2021
Later among the works it cites.
The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks
Emmanuel Abbe, Enric Boix Adsera, and Theodor Misiakiewicz · 2022
Later among the works it cites.
A random matrix perspective on mixtures of nonlinearities in high dimensions
Ben Adlam, Jake A Levinson, and Jeffrey Pennington · 2022
Later among the works it cites.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
Jimmy Ba, Murat A Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, and Greg Yang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science
Roman Vershynin · 2018
Cited alongside, same era.
Universality in learning from linear measurements
Ehsan Abbasi, Fariborz Salehi, and Babak Hassibi · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Cited alongside, same era.
The matrix Dyson equation and its applications for random matrices
Laszlo Erdös · 2019
Cited alongside, same era.
The spectral norm of random inner-product kernel matrices
Zhou Fan and Andrea Montanari · 2019
Cited alongside, same era.
Learning neural networks with two nonlinear layers in polynomial time
Surbhi Goel and Adam R Klivans · 2019
Cited alongside, same era.
Lucas Benigni and Sandrine Péché · 2022
Later among the works it cites.
Hardness of noise-free learning for two-hidden-layer neural networks
Sitan Chen, Aravind Gollakota, Adam Klivans, and Raghu Meka · 2022
Later among the works it cites.
Random Matrix Methods for Machine Learning
Romain Couillet and Zhenyu Liao · 2022
Later among the works it cites.
Neural networks can learn representations with gradient descent
Alex Damian, Jason Lee, and Mahdi Soltanolkotabi · 2022
Later among the works it cites.
The Gaussian equivalence of generative models for learning with shallow neural networks
Sebastian Goldt, Bruno Loureiro, Galen Reeves, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2022
Later among the works it cites.
Hamed Hassani and Adel Javanmard · 2022
Later among the works it cites.
Moving beyond sub-gaussianity in high-dimensional statistics: Applications in covariance estimation and linear regression
Arun Kumar Kuchibhotla and Abhishek Chakrabortty · 2022
Later among the works it cites.
An equivalence principle for the spectrum of random inner-product kernel matrices
Yue M Lu and Horng-Tzer Yau · 2022
Later among the works it cites.
Theodor Misiakiewicz · 2022
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari · 2022
Later among the works it cites.
Universality of empirical risk minimization
Andrea Montanari and Basil N Saeed · 2022
Later among the works it cites.
Feature learning in neural networks and kernel machines that recursively learn features
Adityanarayanan Radhakrishnan, Daniel Beaglehole, Parthe Pandit, and Mikhail Belkin · 2022
Later among the works it cites.
Trainability and accuracy of artificial neural networks: An interacting particle system approach
Grant Rotskoff and Eric Vanden-Eijnden · 2022
Later among the works it cites.
A theoretical analysis on feature learning in neural networks: Emergence from inputs and advantage over fixed features
Zhenmei Shi, Junyi Wei, and Yingyu Liang · 2022
Later among the works it cites.
Spectral evolution and invariance in linear-width neural networks
Zhichao Wang, Andrew Engel, Anand Sarwate, Ioana Dumitriu, and Tony Chiang · 2022
Later among the works it cites.
Precise learning curves and higher-order scalings for dot-product kernel regression
Lechao Xiao, Hong Hu, Theodor Misiakiewicz, Yue Lu, and Jeffrey Pennington · 2022
Later among the works it cites.
A theoretical analysis on feature learning in neural networks: Emergence from inputs and advantage over fixed features
Shi Zhenmei, Junyi Wei, and Yingyu Liang · 2022
Later among the works it cites.
Beyond the universal law of robustness: Sharper laws for random features and neural tangent kernels
Simone Bombari, Shayan Kiyani, and Marco Mondelli · 2023
Closest in time.
Stability, generalization and privacy: Precise analysis for random and NTK features
Simone Bombari and Marco Mondelli · 2023
Closest in time.
Learning time-scales in two-layers neural networks
Raphaël Berthier, Andrea Montanari, and Kangjie Zhou · 2023
Closest in time.
Precise asymptotic analysis of deep random feature models
David Bosch, Ashkan Panahi, and Babak Hassibi · 2023
Closest in time.
Provable multi-task representation learning by two-layer relu neural networks
Liam Collins, Hamed Hassani, Mahdi Soltanolkotabi, Aryan Mokhtari, and Sanjay Shakkottai · 2023
Closest in time.
Bayes-optimal learning of deep random networks of extensive-width
Hugo Cui, Florent Krzakala, and Lenka Zdeborova · 2023
Closest in time.
On double-descent in uncertainty quantification in overparametrized models
Lucas Clarté, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2023
Closest in time.
Learning two-layer neural networks, one (giant) step at a time
Yatin Dandi, Florent Krzakala, Bruno Loureiro, Luca Pesce, and Ludovic Stephan · 2023
Closest in time.
Universality laws for high-dimensional learning with random features
Hong Hu and Yue M Lu · 2023
Closest in time.
Demystifying disagreement-on-the-line in high dimensions
Donghwan Lee, Behrad Moniri, Xinmeng Huang, Edgar Dobriban, and Hamed Hassani · 2023
Closest in time.
Provable guarantees for nonlinear feature learning in three-layer neural networks
Eshaan Nichani, Alex Damian, and Jason D Lee · 2023
Closest in time.
Some notes on concentration for
Holger Sambale · 2023
Closest in time.
Separation of scales and a thermodynamic description of feature learning in some cnns
Inbar Seroussi, Gadi Naveh, and Zohar Ringel · 2023
Closest in time.
Learning hierarchical polynomials with three-layer neural networks
Zihao Wang, Eshaan Nichani, and Jason D Lee · 2023
Closest in time.