Fetching the paper…
Reading the bibliography…
An infinitely wide model is a weighted integration $\int \varphi(x,v) d \mu(v)$ of feature maps.
Universal approximation bounds for superpositions of a sigmoidal function
A. R. Barron · 1993
Earlier work this paper cites.
Constructive Approximation
R. A. DeVore and G. G. Lorentz · 1993
Earlier work this paper cites.
An integral representation of functions using three-layered betworks and their approximation bounds
N. Murata · 1996
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Radford M. Neal · 1996
Earlier work this paper cites.
Weak Convergence and Empirical Processes: With Applications to Statistics
A. van der Vaart and J. Wellner · 1996
Earlier work this paper cites.
Integral Probability Metrics and Their Generating Classes of Functions
A. Müller · 1997
Earlier work this paper cites.
Ridgelets: theory and applications
E. J. Candès · 1998
Earlier work this paper cites.
Harmonic analysis of neural networks
E. J. Candès · 1999
Earlier work this paper cites.
Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond
B. Schölkopf and A. Smola · 2001
Earlier work this paper cites.
Rademacher and Gaussian Complexities: Risk Bounds and Structural Results
P. L. Bartlett and S. Mendelson · 2002
Earlier work this paper cites.
Improving the sample complexity using global data
S. Mendelson · 2002
Earlier work this paper cites.
A conditional gradient method with linear rate of convergence for solving convex linear systems
A. Beck and M. Teboulle · 2004
Earlier work this paper cites.
Convex neural networks
Y. Bengio, N. Le Roux, P. Vincent, O. Delalleau, and P. Marcotte · 2006
Earlier work this paper cites.
Continuous Neural Networks
N. Le Roux and Y. Bengio · 2007
Earlier work this paper cites.
Random Features for Large-Scale Kernel Machines
A. Rahimi and B. Recht · 2008
Cited alongside, same era.
Support Vector Machines
I. Steinwart and A. Christmann · 2008
Cited alongside, same era.
Weighted Sums of Random Kitchen Sinks: Replacing minimization with randomization in learning
A. Rahimi and B. Recht · 2009
Cited alongside, same era.
Super-Samples from Kernel Herding
Y. Chen, M. Welling, and A. Smola · 2010
Cited alongside, same era.
Complexity estimates based on integral transforms induced by computational units
V. Kůrková · 2012
Cited alongside, same era.
Boosting: Foundations and Algorithms
R. E. Schapire and Y. Freund · 2012
Cited alongside, same era.
Double Continuum Limit of Deep Neural Networks
S. Sonoda and N. Murata · 2017
Later among the works it cites.
Embedding of R C D ∗ ( K , N ) RCD^{*}(K,N) spaces in L 2 L^{2} via eigenfunctions
L. Ambrosio, S. Honda, J. W. Portegies, and D. Tewodrose · 2018
Later among the works it cites.
Stein Points
W. Y. Chen, L. Mackey, J. Gorham, F.-X. Briol, and C. J. Oates · 2018
Later among the works it cites.
On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport
L. Chizat and F. Bach · 2018
Later among the works it cites.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Later among the works it cites.
A mean field view of the landscape of two-layer neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Jaggi · 2013
Cited alongside, same era.
Sampling hidden parameters from oracle distribution
S. Sonoda and N. Murata · 2014
Cited alongside, same era.
Mathematical Foundations of Infinite-Dimensional Statistical Models
E. Giné and R. Nickl · 2015
Cited alongside, same era.
On the Global Linear Convergence of Frank-Wolfe Optimization Variants
S. Lacoste-Julien and M. Jaggi · 2015
Cited alongside, same era.
Probabilistic Integration: A Role for Statisticians in Numerical Analysis?
F.-X. Briol, C. J. Oates, M. Girolami, M. A. Osborne, and D. Sejdinovic · 2016
Cited alongside, same era.
Simulation and the Monte Carlo Method
R. Y. Rubinstein and D. P. Kroese · 2016
Cited alongside, same era.
S. Mei, A. Montanari, and P.-M. Nguyen · 2018
Later among the works it cites.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks
G. Rotskoff and E. Vanden-Eijnden · 2018
Later among the works it cites.
Kernel Distribution Embeddings: Universal Kernels, Characteristic Kernels and Kernel Metrics on Distributions
C.-J. Simon-Gabriel and B. Schölkopf · 2018
Later among the works it cites.
Fast generalization error bound of deep learning from a kernel perspective
T. Suzuki · 2018
Later among the works it cites.
High-Dimensional Probability: An Introduction with Applications in Data Science
R. Vershynin · 2018
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
M. Belkin, D. Hsu, S. Ma, and S. Mandal · 2019
Closest in time.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
J. Frankle and M. Carbin · 2019
Closest in time.
Surprises in High-Dimensional Ridgeless Least Squares Interpolation
T. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani · 2019
Closest in time.
Mean Field Analysis of Neural Networks: A Law of Large Numbers
J. Sirignano and K. Spiliopoulos · 2020
Closest in time.