Fetching the paper…
Reading the bibliography…
Consider the problem: given the data pair $(\mathbf{x}, \mathbf{y})$ drawn from a population with $f_*(x) = \mathbf{E}[\mathbf{y} | \mathbf{x} = x]$, specify a neural network model and run gradient flow on the weights over time until reaching any stationarity.
Optimal rates of convergence for nonparametric estimators
Charles J Stone · 1980
Earlier work this paper cites.
Nonparametric maximum likelihood estimation by the method of sieves
Stuart Geman and Chii-Ruey Hwang · 1982
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Spline models for observational data , volume 59
Grace Wahba · 1990
Earlier work this paper cites.
Universal approximation using radial-basis-function networks
Jooyoung Park and Irwin W Sandberg · 1991
Earlier work this paper cites.
A simple lemma on greedy approximation in hilbert space and convergence rates for projection pursuit regression and neural network training
Lee K Jones · 1992
Earlier work this paper cites.
On the relationship between generalization error, hypothesis complexity, and sample complexity for radial basis functions
Partha Niyogi and Federico Girosi · 1996
Earlier work this paper cites.
The variational formulation of the fokker–planck equation
Richard Jordan, David Kinderlehrer, and Felix Otto · 1998
Earlier work this paper cites.
Statistical learning theory. 1998 , volume 3
Vladimir Vapnik · 1998
Earlier work this paper cites.
Convex neural networks
Yoshua Bengio, Nicolas L Roux, Pascal Vincent, Olivier Delalleau, and Patrice Marcotte · 2006
Earlier work this paper cites.
Approximation and learning by greedy algorithms
Andrew R Barron, Albert Cohen, Wolfgang Dahmen, Ronald A DeVore, et al · 2008
Earlier work this paper cites.
Risk of penalized least squares, greedy selection and l1-penalization for flexible function libraries
Cong Huang, Gerald HL Cheang, and Andrew R Barron · 2008
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Cited alongside, same era.
Neural network learning: Theoretical foundations
Martin Anthony and Peter L Bartlett · 2009
Cited alongside, same era.
Kernel methods for deep learning
Youngmin Cho and Lawrence K Saul · 2009
Cited alongside, same era.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Ali Rahimi and Benjamin Recht · 2009
Cited alongside, same era.
Essays in analysis
Bill Casselman · 2014
Cited alongside, same era.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Later among the works it cites.
Max H Farrell, Tengyuan Liang, and Sanjog Misra · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Representational power of relu networks and polynomial kernels: Beyond worst-case analysis
Frederic Koehler and Andrej Risteski · 2018
Later among the works it cites.
Just interpolate: Kernel” ridgeless” regression can generalize
Tengyuan Liang and Alexander Rakhlin · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Cited alongside, same era.
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2017
Cited alongside, same era.
Why and when can deep-but not shallow-networks avoid the curse of dimensionality: a review
Tomaso Poggio, Hrushikesh Mhaskar, Lorenzo Rosasco, Brando Miranda, and Qianli Liao · 2017
Cited alongside, same era.
Generalization properties of learning with random features
Alessandro Rudi and Lorenzo Rosasco · 2017
Cited alongside, same era.
A note on lazy training in supervised differentiable programming
Lenaic Chizat and Francis Bach · 2018
Cited alongside, same era.
Later among the works it cites.
Gradient descent quantizes relu network features
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Later among the works it cites.
A mean field view of the landscape of two-layers neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Later among the works it cites.
Consistency of interpolation with laplace kernels is a high-dimensional phenomenon
Alexander Rakhlin and Xiyu Zhai · 2018
Later among the works it cites.
Grant M Rotskoff and Eric Vanden-Eijnden · 2018
Later among the works it cites.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Closest in time.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Closest in time.
Mean field analysis of neural networks: A central limit theorem
Justin Sirignano and Konstantinos Spiliopoulos · 2019
Closest in time.