Fetching the paper…
Reading the bibliography…
It is currently known how to characterize functions that neural networks can learn with SGD for two extremal parameterizations: neural networks in the linear regime, and neural networks with no structural constraints.
Waring’s problem for the ring of polynomials
DJ Newman and Morton Slater · 1979
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R Barron · 1993
Earlier work this paper cites.
Weakly learning DNF and characterizing statistical query learning using fourier analysis
Avrim Blum, Merrick L. Furst, Jeffrey C. Jackson, Michael J. Kearns, Yishay Mansour, and Steven Rudich · 1994
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Michael Kearns · 1998
Earlier work this paper cites.
Distributional and lq norm inequalities for polynomials over convex bodies in rn
Anthony Carbery and James Wright · 2001
Earlier work this paper cites.
Noise-tolerant learning, the parity problem, and the statistical query model
Avrim Blum, Adam Kalai, and Hal Wasserman · 2003
Earlier work this paper cites.
New lower bounds for statistical query learning
Ke Yang · 2005
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate o (1/n)
Francis Bach and Eric Moulines · 2013
Earlier work this paper cites.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Ohad Shamir and Tong Zhang · 2013
Earlier work this paper cites.
Memory, communication, and statistical queries
Jacob Steinhardt, Gregory Valiant, and Stefan Wager · 2015
Earlier work this paper cites.
A general characterization of the statistical query complexity
Vitaly Feldman · 2016
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Earlier work this paper cites.
Statistical algorithms and a lower bound for detecting planted cliques
Vitaly Feldman, Elena Grigorescu, Lev Reyzin, Santosh S. Vempala, and Ying Xiao · 2017
Earlier work this paper cites.
Statistical query algorithms for mean vector estimation and stochastic convex optimization
Vitaly Feldman, Cristóbal Guzmán, and Santosh Vempala · 2017
Earlier work this paper cites.
Prateek Jain, Sham M Kakade, Rahul Kidambi, Praneeth Netrapalli, Venkata Krishna Pillutla, and Aaron Sidford · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaïc Chizat and Francis Bach · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Parallelizing stochastic gradient descent for least squares regression: mini-batching, averaging, and model misspecification
Prateek Jain, Sham Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Cited alongside, same era.
Neural networks as interacting particle systems: Asymptotic convexity of the loss landscape and universal scaling of the approximation error
Grant M Rotskoff and Eric Vanden-Eijnden · 2018
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Cited alongside, same era.
What can resnet learn efficiently, going beyond kernels?
Zeyuan Allen-Zhu and Yuanzhi Li · 2019
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Cited alongside, same era.
Learning over-parametrized two-layer neural networks beyond NTK
Yuanzhi Li, Tengyu Ma, and Hongyang R. Zhang · 2020
Later among the works it cites.
The implications of local correlation on learning some deep functions
Eran Malach and Shai Shalev-Shwartz · 2020
Later among the works it cites.
A rigorous framework for the mean field limit of multilayer neural networks
Phan-Minh Nguyen and Huy Tuan Pham · 2020
Later among the works it cites.
Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
Taiji Suzuki and Shunta Akiyama · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
The convergence rate of neural networks for learned functions of different frequencies
Ronen Basri, David Jacobs, Yoni Kasten, and Shira Kritchman · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
Limitations of lazy training of two-layers neural networks
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Cited alongside, same era.
Making the last iterate of sgd information theoretically optimal
Prateek Jain, Dheeraj Nagaraj, and Praneeth Netrapalli · 2019
Cited alongside, same era.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Cited alongside, same era.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Cited alongside, same era.
Nonparametric regression using deep neural networks with relu activation function
Johannes Schmidt-Hieber · 2020
Later among the works it cites.
Mean field analysis of neural networks: A central limit theorem
Justin Sirignano and Konstantinos Spiliopoulos · 2020
Later among the works it cites.
On the convergence of gradient descent training for two-layer relu-networks in the mean field regime
Stephan Wojtowytsch · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Later among the works it cites.
The staircase property: How hierarchical structure can guide deep learning
Emmanuel Abbe, Enric Boix-Adserà, Matthew S. Brennan, Guy Bresler, and Dheeraj Nagaraj · 2021
Later among the works it cites.
On the power of differentiable learning versus PAC and SQ learning
Emmanuel Abbe, Pritish Kamath, Eran Malach, Colin Sandon, and Nathan Srebro · 2021
Later among the works it cites.
Towards understanding the spectral bias of deep learning
Yuan Cao, Zhiying Fang, Yue Wu, Ding-Xuan Zhou, and Quanquan Gu · 2021
Later among the works it cites.
Minimum complexity interpolation in random features models
Michael Celentano, Theodor Misiakiewicz, and Andrea Montanari · 2021
Later among the works it cites.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
Later among the works it cites.
When do neural networks outperform kernel methods?
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
Later among the works it cites.
On the approximation power of two-layer networks of random relus
Daniel Hsu, Clayton Sanford, Rocco A Servedio, and Emmanouil-Vasileios Vlatakis-Gkaragkounis · 2021
Later among the works it cites.
Local signal adaptivity: Provable feature learning in neural networks beyond kernels
Stefani Karp, Ezra Winston, Yuanzhi Li, and Aarti Singh · 2021
Later among the works it cites.
Quantifying the benefit of using differentiable learning over tangent kernels
Eran Malach, Pritish Kamath, Emmanuel Abbe, and Nathan Srebro · 2021
Later among the works it cites.
Generalization error of random feature and kernel methods: Hypercontractivity and kernel matrix concentration
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
Later among the works it cites.
Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed
Maria Refinetti, Sebastian Goldt, Florent Krzakala, and Lenka Zdeborova · 2021
Later among the works it cites.