Fetching the paper…
Reading the bibliography…
Significant theoretical work has established that in specific regimes, neural networks trained by gradient descent behave like kernel methods.
S. Arora, S. S. Du, W. Hu, Z. Li, and R. Wang · 1901
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
S. Arora, S. S. Du, W. Hu, Z. Li, R. Salakhutdinov, and R. Wang · 1904
Earlier work this paper cites.
Linearized two-layers neural networks in high dimension
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari · 1904
Earlier work this paper cites.
On the power and limitations of random features for understanding neural networks
G. Yehudai and O. Shamir · 1904
Earlier work this paper cites.
On the power and limitations of random features for understanding neural networks
G. Yehudai and O. Shamir · 1904
Earlier work this paper cites.
Towards understanding hierarchical learning: Benefits of neural representations
M. Chen, Y. Bai, J. Lee, T. Zhao, H. Wang, C. Xiong, and R. Socher · 2006
Earlier work this paper cites.
Wick powers in stochastic pdes: an introduction
G. D. Prato and L. Tubaro · 2007
Earlier work this paper cites.
Characterizing statistical query learning: Simplified notions and proofs
B. Szörényi · 2009
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Earlier work this paper cites.
Reliably learning the relu in polynomial time
S. Goel, V. Kanade, A. Klivans, and J. Thaler · 2017
Earlier work this paper cites.
Learning relus via gradient descent
M. Soltanolkotabi · 2017
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
L. Chizat and F. Bach · 2018
Earlier work this paper cites.
On the power of over-parametrization in neural networks with quadratic activation
S. Du and J. Lee · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Y. Li and Y. Liang · 2018
Earlier work this paper cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Y. Li, T. Ma, and H. Zhang · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
S. Mei, A. Montanari, and P.-M. Nguyen · 2018
Earlier work this paper cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
M. Soltanolkotabi, A. Javanmard, and J. D. Lee · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science , volume 47
R. Vershynin · 2018
Cited alongside, same era.
Stochastic gradient descent optimizes over-parameterized deep relu networks
D. Zou, Y. Cao, D. Zhou, and Q. Gu · 2018
Cited alongside, same era.
What can resnet learn efficiently, going beyond kernels?
Z. Allen-Zhu and Y. Li · 2019
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Z. Allen-Zhu, Y. Li, and Y. Liang · 2019
Cited alongside, same era.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Y. Bai and J. D. Lee · 2020
Later among the works it cites.
Taylorized training: Towards better approximation of neural network training at finite width
Y. Bai, B. Krause, H. Wang, C. Xiong, and R. Socher · 2020
Later among the works it cites.
Learning polynomials of few relevant dimensions
S. Chen and R. Meka · 2020
Later among the works it cites.
Learning parities with neural networks
A. Daniely and E. Malach · 2020
Later among the works it cites.
Algorithms and sq lower bounds for pac learning one-hidden-layer relu networks
I. Diakonikolas, D. M. Kane, V. Kontonis, and N. Zarifis · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Cao and Q. Gu · 2019
Cited alongside, same era.
On lazy training in differentiable programming
L. Chizat, E. Oyallon, and F. Bach · 2019
Cited alongside, same era.
Asymptotics of wide networks from feynman diagrams
E. Dyer and G. Gur-Ari · 2019
Cited alongside, same era.
Time/accuracy tradeoffs for learning a relu with respect to gaussian marginals
S. Goel, S. Karmalkar, and A. Klivans · 2019
Cited alongside, same era.
Concentration inequalities for polynomials in alpha-sub-exponential random variables
F. Gotze, H. Sambale, and A. Sinulis · 2019
Cited alongside, same era.
Dynamics of deep neural networks and neural tangent hierarchy
J. Huang and H.-T. Yau · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
J. Lee, L. Xiao, S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington · 2019
Cited alongside, same era.
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari · 2020
Later among the works it cites.
Superpolynomial lower bounds for learning one-layer neural networks using gradient descent
S. Goel, A. Gollakota, Z. Jin, S. Karmalkar, and A. Klivans · 2020
Later among the works it cites.
Analysis of a two-layer neural network via displacement convexity
A. Javanmard, M. Mondelli, A. Montanari, et al · 2020
Later among the works it cites.
Learning over-parametrized two-layer relu neural networks beyond ntk
Y. Li, T. Ma, and H. R. Zhang · 2020
Later among the works it cites.
Bad global minima exist and sgd can reach them
S. Liu, D. Papailiopoulos, and D. Achlioptas · 2020
Later among the works it cites.
Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
S. Oymak and M. Soltanolkotabi · 2020
Later among the works it cites.
Mean field analysis of neural networks: a central limit theorem
J. Sirignano and K. Spiliopoulos · 2020
Later among the works it cites.
Beyond lazy training for over-parameterized tensor decomposition
X. Wang, C. Wu, J. D. Lee, T. Ma, and R. Ge · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
B. Woodworth, S. Gunasekar, J. D. Lee, E. Moroshko, P. Savarese, I. Golan, D. Soudry, and N. Srebro · 2020
Later among the works it cites.
Generalization guarantees for neural architecture search with train-validation split
S. Oymak, M. Li, and M. Soltanolkotabi · 2021
Later among the works it cites.
Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction
D. Stöger and M. Soltanolkotabi · 2021
Later among the works it cites.
E. Abbe, E. Boix-Adsera, and T. Misiakiewicz · 2022
Closest in time.