Fetching the paper…
Reading the bibliography…
A recent series of theoretical works showed that the dynamics of neural networks with a certain initialisation are well-captured by kernel methods.
Three unfinished works on the optimal storage capacity of networks
Gardner, E. and Derrida, B · 1989
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Barron, A. R · 1993
Earlier work this paper cites.
Learning by on-line gradient descent
Biehl, M. and Schwarze, H · 1995
Earlier work this paper cites.
On-line backpropagation in two-layered neural networks
Riegler, P. and Biehl, M · 1995
Earlier work this paper cites.
Transient dynamics of on-line learning in two-layered neural networks
Biehl, M., Riegler, P., and Wöhler, C · 1996
Earlier work this paper cites.
Learning with Noise and Regularizers Multilayer Neural Networks
Saad, D. and Solla, S · 1997
Earlier work this paper cites.
The gaussian equivalence of generative models for learning with two-layer neural networks
Goldt, S., Loureiro, B., Reeves, G., Mézard, M., Krzakala, F., and Zdeborová, L · 2006
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Caponnetto, A. and De Vito, E · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B · 2008
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Rahimi, A. and Recht, B · 2009
Earlier work this paper cites.
On-line learning in neural networks , volume 17
Saad, D · 2009
Earlier work this paper cites.
Optimal rates for regularized least squares regression
Steinwart, I., Hush, D. R., Scovel, C., et al · 2009
Earlier work this paper cites.
The spectrum of kernel random matrices
El Karoui, N · 2010
Earlier work this paper cites.
On the computational efficiency of training neural networks
Livni, R., Shalev-Shwartz, S., and Shamir, O · 2014
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Bach, F · 2017
Earlier work this paper cites.
Nonlinear random matrix theory for deep learning
Pennington, J. and Worah, P · 2017
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2018
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Arora, S., Cohen, N., and Hazan, E · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L. and Bach, F · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Deep neural networks as gaussian processes
Lee, J., Sohl-Dickstein, J., Pennington, J., Novak, R., Schoenholz, S., and Bahri, Y · 2018
Earlier work this paper cites.
Learning Overparameterized Neural Networks via Stochastic Gradient Descent on Structured Data
Li, Y. and Liang, Y · 2018
Cited alongside, same era.
On the spectrum of random features maps of high dimensional data
Liao, Z. and Couillet, R · 2018
Cited alongside, same era.
A random matrix approach to neural networks
Louart, C., Liao, Z., and Couillet, R · 2018
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
Matthews, A. G. d. G., Hron, J., Rowland, M., Turner, R., and Ghahramani, Z · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Mei, S., Montanari, A., and Nguyen, P · 2018
Cited alongside, same era.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks
Rotskoff, G. and Vanden-Eijnden, E · 2018
High dimensional classification via empirical risk minimization: Improvements and optimality
Mai, X. and Liao, Z · 2019
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Mei, S. and Montanari, A · 2019
Later among the works it cites.
Mean field analysis of neural networks: A central limit theorem
Sirignano, J. and Spiliopoulos, K · 2019
Later among the works it cites.
A solvable high-dimensional model of gan
Wang, C., Hu, H., and Lu, Y · 2019
Later among the works it cites.
Regularization matters: Generalization and optimization of neural nets v.s. their induced kernel
Wei, C., Lee, J. D., Liu, Q., and Ma, T · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Spurious local minima are common in two-layer relu neural networks
Safran, I. and Shamir, O · 2018
Cited alongside, same era.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Scholkopf, B. and Smola, A · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Arora, S., Du, S., Hu, W., Li, Z., Salakhutdinov, R., and Wang, R · 2019
Cited alongside, same era.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y. and Gu, Q · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Cited alongside, same era.
On the power and limitations of random features for understanding neural networks
Yehudai, G. and Shamir, O · 2019
Later among the works it cites.
Data-dependence of plateau phenomenon in learning with neural network — statistical mechanical analysis
Yoshida, Y. and Okada, M · 2019
Later among the works it cites.
Statistical mechanical analysis of learning dynamics of two-layer perceptron with multiple output units
Yoshida, Y., Karakida, R., Okada, M., and Amari, S.-I · 2019
Later among the works it cites.
A classification for the performance of online sgd for high-dimensional inference
Arous, G. B., Gheissari, R., and Jagannath, A · 2020
Later among the works it cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Chizat, L. and Bach, F · 2020
Later among the works it cites.
Separability and geometry of object manifolds in deep neural networks
Cohen, U., Chung, S., Lee, D., and Sompolinsky, H · 2020
Later among the works it cites.
Learning parities with neural networks
Daniely, A. and Malach, E · 2020
Later among the works it cites.
Disentangling feature and lazy training in deep neural networks
Geiger, M., Spigler, S., Jacot, A., and Wyart, M · 2020
Later among the works it cites.
When do neural networks outperform kernel methods?
Ghorbani, B., Mei, S., Misiakiewicz, T., and Montanari, A · 2020
Later among the works it cites.
Analytic study of double descent in binary classification: The impact of loss
Kini, G. R. and Thrampoulidis, C · 2020
Later among the works it cites.
Learning over-parametrized two-layer neural networks beyond ntk
Li, Y., Ma, T., and Zhang, H. R · 2020
Later among the works it cites.
Neural kernels without tangents
Shankar, V., Fang, A., Guo, W., Fridovich-Keil, S., Ragan-Kelley, J., Schmidt, L., and Recht, B · 2020
Later among the works it cites.
Suzuki, T. and Akiyama, S · 2020
Later among the works it cites.
Mei, S., Misiakiewicz, T., and Montanari, A · 2021
Closest in time.
Geometric compression of invariant manifolds in neural networks
Paccolat, J., Petrini, L., Geiger, M., Tyloo, K., and Wyart, M · 2021
Closest in time.