Fetching the paper…
Reading the bibliography…
For high-dimensional Gaussian data, we investigate theoretically how the features of a two-layer neural network adapt to the structure of the target function through a few large batch gradient descent steps, leading to an improvement in the approximation capacity from initialization.
Note on N-dimensional hermite polynomials
Harold Grad · 1949
Earlier work this paper cites.
Probability in Banach Spaces: Isoperimetry and Processes
Michel Ledoux and Michel Talagrand · 1991
Earlier work this paper cites.
Decoupling Inequalities for the Tail Probabilities of Multivariate $U$-Statistics
Victor H. de la Pena and S. J. Montgomery-Smith · 1995
Earlier work this paper cites.
On-line learning in soft committee machines
David Saad and Sara A. Solla · 1995
Earlier work this paper cites.
Weak Convergence and Empirical Processes: With Applications to Statistics
Aad van der Vaart and Jon Wellner · 1996
Earlier work this paper cites.
Gaussian hilbert spaces
Svante Janson · 1997
Earlier work this paper cites.
Statistical mechanics of support vector networks
Rainer Dietrich, Manfred Opper, and Haim Sompolinsky · 1999
Earlier work this paper cites.
A multilinear singular value decomposition
Lieven De Lathauwer, Bart De Moor, and Joos Vandewalle · 2000
Earlier work this paper cites.
Universal learning curves of support vector machines
M. Opper and R. Urbanczik · 2001
Earlier work this paper cites.
Feynman diagrams for pedestrians and mathematicians
Michael Polyak · 2005
Earlier work this paper cites.
Randomized gossip algorithms
S. Boyd, A. Ghosh, B. Prabhakar, and D. Shah · 2006
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
Multilinear Algebra
W.H. Greub · 2012
Earlier work this paper cites.
A probabilistic theory of pattern recognition , volume 31
Luc Devroye, László Györfi, and Gábor Lugosi · 2013
Earlier work this paper cites.
Probability theory: a comprehensive course
Achim Klenke · 2013
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Generalization properties of learning with random features
Alessandro Rudi and Lorenzo Rosasco · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Earlier work this paper cites.
Learning single-index models in gaussian space
Rishabh Dudeja and Daniel Hsu · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science , volume 47
Roman Vershynin · 2018
Earlier work this paper cites.
On Lazy Training in Differentiable Programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Earlier work this paper cites.
Limitations of lazy training of two-layers neural network
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Earlier work this paper cites.
Concentration inequalities for polynomials in α \alpha -sub-exponential random variables
Friedrich Gotze, Holger Sambale, and Arthur Sinulis · 2019
Cited alongside, same era.
Sgd on neural networks learns functions of increasing complexity
Dimitris Kalimeris, Gal Kaplun, Preetum Nakkiran, Benjamin Edelman, Tristan Yang, Boaz Barak, and Haofeng Zhang · 2019
Cited alongside, same era.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Cited alongside, same era.
Spectrum dependent learning curves in kernel regression and wide neural networks
Blake Bordelon, Abdulkadir Canatar, and Cengiz Pehlevan · 2020
Cited alongside, same era.
A precise performance analysis of learning with random features, 2020
Oussama Dhifallah and Yue M. Lu · 2020
Cited alongside, same era.
Learning single-index models with shallow neural networks
Alberto Bietti, Joan Bruna, Clayton Sanford, and Min Jae Song · 2022
Later among the works it cites.
Gradient flow dynamics of shallow relu networks for square loss and orthogonal inputs
Etienne Boursier, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2022
Later among the works it cites.
Error rates for kernel classification under source and capacity conditions, 2022
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2022
Later among the works it cites.
Neural networks can learn representations with gradient descent
Alexandru Damian, Jason Lee, and Mahdi Soltanolkotabi · 2022
Later among the works it cites.
The gaussian equivalence of generative models for learning with shallow neural networks
Sebastian Goldt, Bruno Loureiro, Galen Reeves, Florent Krzakala, Marc Mezard, and Lenka Zdeborova · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mezard, and Lenka Zdeborova · 2020
Cited alongside, same era.
When do neural networks outperform kernel methods?
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2020
Cited alongside, same era.
A review of applications in federated learning
Li Li, Yuxi Fan, Mike Tse, and Kuo-Yi Lin · 2020
Cited alongside, same era.
The zero set of a real analytic function
Boris Samuilovich Mityagin · 2020
Cited alongside, same era.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2020
Cited alongside, same era.
Understanding deep learning is also a job for physicists
Lenka Zdeborová · 2020
Cited alongside, same era.
The staircase property: How hierarchical structure can guide deep learning
Emmanuel Abbe, Enric Boix-Adsera, Matthew S Brennan, Guy Bresler, and Dheeraj Nagaraj · 2021
Cited alongside, same era.
Hong Hu and Yue M Lu · 2022
Later among the works it cites.
Fluctuations, bias, variance &; ensemble of learners: Exact asymptotics for convex losses in high-dimension
Bruno Loureiro, Cedric Gerbelot, Maria Refinetti, Gabriele Sicuro, and Florent Krzakala · 2022
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari · 2022
Later among the works it cites.
Universality of empirical risk minimization
Andrea Montanari and Basil N. Saeed · 2022
Later among the works it cites.
Learning sparse features can lead to overfitting in neural networks, 2022
Leonardo Petrini, Francesco Cagnetta, Eric Vanden-Eijnden, and Matthieu Wyart · 2022
Later among the works it cites.
Trainability and accuracy of artificial neural networks: An interacting particle system approach
Grant Rotskoff and Eric Vanden-Eijnden · 2022
Later among the works it cites.
The eigenlearning framework: A conservation law perspective on kernel regression and wide neural networks, 2022
James B. Simon, Madeline Dickens, Dhruva Karkada, and Michael R. DeWeese · 2022
Later among the works it cites.
Precise learning curves and higher-order scalings for dot-product kernel regression
Lechao Xiao, Hong Hu, Theodor Misiakiewicz, Yue Lu, and Jeffrey Pennington · 2022
Later among the works it cites.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics, 2023
Emmanuel Abbe, Enric Boix-Adsera, and Theodor Misiakiewicz · 2023
Closest in time.
Luca Arnaboldi, Ludovic Stephan, Florent Krzakala, and Bruno Loureiro · 2023
Closest in time.
Learning time-scales in two-layers neural networks, 2023
Raphaël Berthier, Andrea Montanari, and Kangjie Zhou · 2023
Closest in time.
Dynamics of finite width kernel and prediction fluctuations in mean field neural networks, 2023
Blake Bordelon and Cengiz Pehlevan · 2023
Closest in time.
Precise asymptotic analysis of deep random feature models, 2023
David Bosch, Ashkan Panahi, and Babak Hassibi · 2023
Closest in time.
Alex Damian, Eshaan Nichani, Rong Ge, and Jason D. Lee · 2023
Closest in time.
Universality laws for gaussian mixtures in generalized linear models, 2023
Yatin Dandi, Ludovic Stephan, Florent Krzakala, Bruno Loureiro, and Lenka Zdeborová · 2023
Closest in time.
Deterministic equivalent and error universality of deep random features learning, 2023
Dominik Schröder, Hugo Cui, Daniil Dmitriev, and Bruno Loureiro · 2023
Closest in time.
Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks
Abdulkadir Canatar, Blake Bordelon, and Cengiz Pehlevan · 2041
Closest in time.
Separation of scales and a thermodynamic description of feature learning in some cnns
Inbar Seroussi, Gadi Naveh, and Zohar Ringel · 2041
Closest in time.