Fetching the paper…
Reading the bibliography…
Recent research shows that the dynamics of an infinitely wide neural network (NN) trained by gradient descent can be characterized by Neural Tangent Kernel (NTK) \citep{jacot2018neural}.
S. Arora, S. S. Du, W. Hu, Z. Li, and R. Wang · 1901
Earlier work this paper cites.
Graph neural tangent kernel: Fusing graph neural networks with graph kernels
S. S. Du, K. Hou, B. Póczos, R. Salakhutdinov, R. Wang, and K. Xu · 1905
Earlier work this paper cites.
Harnessing the power of infinitely wide deep nets on small-data tasks
S. Arora, S. S. Du, Z. Li, R. Salakhutdinov, R. Wang, and D. Yu · 1910
Earlier work this paper cites.
Networks for approximation and learning
T. Poggio and F. Girosi · 1990
Earlier work this paper cites.
A training algorithm for optimal margin classifiers
B. E. Boser, I. M. Guyon, and V. N. Vapnik · 1992
Earlier work this paper cites.
Support-vector networks
C. Cortes and V. Vapnik · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
The connection between regularization operators and support vector kernels
A. J. Smola, B. Schölkopf, and K.-R. Müller · 1998
Earlier work this paper cites.
An introduction to support vector machines and other kernel-based learning methods
N. Cristianini, J. Shawe-Taylor, et al · 2000
Earlier work this paper cites.
The equivalence of support vector machine and regularization neural networks
P. Andras · 2002
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
P. L. Bartlett and S. Mendelson · 2002
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
B. Schölkopf, A. J. Smola, F. Bach, et al · 2002
Earlier work this paper cites.
Subgradient methods
S. Boyd, L. Xiao, and A. Mutapcic · 2003
Earlier work this paper cites.
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
C. Liu, L. Zhu, and M. Belkin · 2003
Earlier work this paper cites.
Kernel methods for pattern analysis
J. Shawe-Taylor, N. Cristianini, et al · 2004
Earlier work this paper cites.
Gradient flows: in metric spaces and in the space of probability measures
L. Ambrosio, N. Gigli, and G. Savaré · 2008
Earlier work this paper cites.
A dual coordinate descent method for large-scale linear svm
C.-J. Hsieh, K.-W. Chang, C.-J. Lin, S. S. Keerthi, and S. Sundararajan · 2008
Earlier work this paper cites.
Support vector machines
I. Steinwart and A. Christmann · 2008
Earlier work this paper cites.
Pegasos: Primal estimated sub-gradient solver for svm
S. Shalev-Shwartz, Y. Singer, N. Srebro, and A. Cotter · 2011
Earlier work this paper cites.
Kernel methods for deep learning
Y. Cho · 2012
Cited alongside, same era.
Deep learning using linear support vector machines
Y. Tang · 2013
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Cited alongside, same era.
Xsede: Accelerating scientific discovery
J. Towns, T. Cockerill, M. Dahan, I. Foster, K. Gaither, A. Grimshaw, V. Hazlewood, S. Lathrop, D. Lifka, G. D. Peterson, R. Roskies, J. R. Scott, and N. Wilkins-Diehr · 2014
Cited alongside, same era.
Convex optimization algorithms
D. P. Bertsekas and A. Scientific · 2015
Cited alongside, same era.
Large-margin softmax loss for convolutional neural networks
W. Liu, Y. Wen, Z. Yu, and M. Yang · 2016
Cited alongside, same era.
Efficient neural network robustness certification with general activation functions
H. Zhang, T.-W. Weng, P.-Y. Chen, C.-J. Hsieh, and L. Daniel · 2018
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2019
Later among the works it cites.
Nearly-tight vc-dimension and pseudodimension bounds for piecewise linear neural networks
P. L. Bartlett, N. Harvey, C. Liaw, and A. Mehrabian · 2019
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Y. Cao and Q. Gu · 2019
Later among the works it cites.
W. Hu, Z. Li, and D. Yu · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the depth of deep neural networks: A theoretical view
S. Sun, W. Chen, L. Wang, X. Liu, and T.-Y. Liu · 2016
Cited alongside, same era.
Deep neural networks as gaussian processes
J. Lee, Y. Bahri, R. Novak, S. S. Schoenholz, J. Pennington, and J. Sohl-Dickstein · 2017
Cited alongside, same era.
Soft-margin softmax for deep classification
X. Liang, X. Wang, Z. Lei, S. Liao, and S. Z. Li · 2017
Cited alongside, same era.
Robust large margin deep neural networks
J. Sokolić, R. Giryes, G. Sapiro, and M. R. Rodrigues · 2017
Cited alongside, same era.
On lazy training in differentiable programming
L. Chizat, E. Oyallon, and F. Bach · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
S. S. Du, X. Zhai, B. Poczos, and A. Singh · 2018
Cited alongside, same era.
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
J. Lee, L. Xiao, S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington · 2019
Later among the works it cites.
Generalization bounds for deep convolutional neural networks
P. M. Long and H. Sedghi · 2019
Later among the works it cites.
Gradient descent maximizes the margin of homogeneous neural networks
K. Lyu and J. Li · 2019
Later among the works it cites.
Lexicographic and depth-sensitive margins in homogeneous and non-homogeneous deep models
M. S. Nacson, S. Gunasekar, J. Lee, N. Srebro, and D. Soudry · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Later among the works it cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
C. Wei, J. Lee, Q. Liu, and T. Ma · 2019
Later among the works it cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
L. Chizat and F. Bach · 2020
Later among the works it cites.
Every model learned by gradient descent is approximately a kernel machine
P. Domingos · 2020
Later among the works it cites.
On the neural tangent kernel of deep networks with orthogonal initialization
W. Huang, W. Du, and R. Y. Da Xu · 2020
Later among the works it cites.
Neural tangents: Fast and easy infinite neural networks in python
R. Novak, L. Xiao, J. Hron, J. Lee, A. A. Alemi, J. Sohl-Dickstein, and S. S. Schoenholz · 2020
Later among the works it cites.
An analytic theory of shallow networks dynamics for hinge loss classification
F. Pellegrini and G. Biroli · 2020
Later among the works it cites.
Feature learning in infinite-width neural networks
G. Yang and E. J. Hu · 2020
Later among the works it cites.
Regularization matters: A nonparametric perspective on overparametrized neural network
T. Hu, W. Wang, C. Lin, and G. Cheng · 2021
Closest in time.
Understanding deep learning (still) requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2021
Closest in time.