Fetching the paper…
Reading the bibliography…
State-of-the-art neural networks are heavily over-parameterized, making the optimization algorithm a crucial ingredient for learning predictive models with good generalization properties.
Multilayer feedforward networks are universal approximators
K. Hornik, M. Stinchcombe, and H. White · 1989
Earlier work this paper cites.
Integral transforms, reproducing kernels and their applications
S. Saitoh · 1997
Earlier work this paper cites.
Approximation theory of the mlp model in neural networks
A. Pinkus · 1999
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
B. Schölkopf and A. J. Smola · 2001
Earlier work this paper cites.
Regularization with dot-product kernels
A. J. Smola, Z. L. Ovari, and R. C. Williamson · 2001
Earlier work this paper cites.
On the mathematical foundations of learning
F. Cucker and S. Smale · 2002
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
A. Caponnetto and E. De Vito · 2007
Earlier work this paper cites.
Training invariant support vector machines using selective sampling
G. Loosli, S. Canu, and L. Bottou · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2008
Earlier work this paper cites.
Kernel methods for deep learning
Y. Cho and L. K. Saul · 2009
Earlier work this paper cites.
Spherical harmonics and approximations on the unit sphere: an introduction
K. Atkinson and W. Han · 2012
Earlier work this paper cites.
Group invariant scattering
S. Mallat · 2012
Earlier work this paper cites.
Invariant scattering convolution networks
J. Bruna and S. Mallat · 2013
Earlier work this paper cites.
Spherical harmonics in p dimensions
C. Efthimiou and C. Frye · 2014
Earlier work this paper cites.
Convolutional kernel networks
J. Mairal, P. Koniusz, Z. Harchaoui, and C. Schmid · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
A. Daniely, R. Frostig, and Y. Singer · 2016
Earlier work this paper cites.
End-to-End Kernel Learning with Supervised Convolutional Kernel Networks
J. Mairal · 2016
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
F. Bach · 2017
Cited alongside, same era.
On the equivalence between kernel quadrature rules and random feature expansions
F. Bach · 2017
Cited alongside, same era.
Sobolev norm learning rates for regularized least-squares algorithm
S. Fischer and I. Steinwart · 2017
Cited alongside, same era.
Generalization properties of learning with random features
A. Rudi and L. Rosasco · 2017
Cited alongside, same era.
Diverse neural network learns true target functions
B. Xie, Y. Liang, and L. Song · 2017
Cited alongside, same era.
To understand deep learning we need to understand kernel learning
M. Belkin, S. Ma, and S. Mandal · 2018
Group invariance, stability to deformations, and complexity of deep convolutional representations
A. Bietti and J. Mairal · 2019
Closest in time.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Y. Cao and Q. Gu · 2019
Closest in time.
On lazy training in differentiable programming
L. Chizat, E. Oyallon, and F. Bach · 2019
Closest in time.
Gradient descent finds global minima of deep neural networks
S. S. Du, J. D. Lee, H. Li, L. Wang, and X. Zhai · 2019
Closest in time.
Gradient descent provably optimizes over-parameterized neural networks
S. S. Du, X. Zhai, B. Poczos, and A. Singh · 2019
Closest in time.
Deep convolutional networks as shallow gaussian processes
A. Garriga-Alonso, L. Aitchison, and C. E. Rasmussen · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
L. Chizat and F. Bach · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Cited alongside, same era.
Deep neural networks as gaussian processes
J. Lee, Y. Bahri, R. Novak, S. S. Schoenholz, J. Pennington, and J. Sohl-Dickstein · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Y. Li and Y. Liang · 2018
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
A. Matthews, M. Rowland, J. Hron, R. E. Turner, and Z. Ghahramani · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
S. Mei, A. Montanari, and P.-M. Nguyen · 2018
Cited alongside, same era.
Linearized two-layers neural networks in high dimension
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
J. Lee, L. Xiao, S. S. Schoenholz, Y. Bahri, J. Sohl-Dickstein, and J. Pennington · 2019
Closest in time.
Just interpolate: Kernel" ridgeless" regression can generalize
T. Liang and A. Rakhlin · 2019
Closest in time.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
S. Mei, T. Misiakiewicz, and A. Montanari · 2019
Closest in time.
Bayesian deep convolutional networks with many channels are gaussian processes
R. Novak, L. Xiao, Y. Bahri, J. Lee, G. Yang, J. Hron, D. A. Abolafia, J. Pennington, and J. Sohl-Dickstein · 2019
Closest in time.
How do infinite width bounded norm networks look in function space?
P. Savarese, I. Evron, D. Soudry, and N. Srebro · 2019
Closest in time.
Gradient dynamics of shallow low-dimensional relu networks
F. Williams, M. Trager, C. Silva, D. Panozzo, D. Zorin, and J. Bruna · 2019
Closest in time.
G. Yang · 2019
Closest in time.
On the power and limitations of random features for understanding neural networks
G. Yehudai and O. Shamir · 2019
Closest in time.
C. Zhang, S. Bengio, and Y. Singer · 2019
Closest in time.
Stochastic gradient descent optimizes over-parameterized deep relu networks
D. Zou, Y. Cao, D. Zhou, and Q. Gu · 2019
Closest in time.