2022

Limitations of the NTK for Understanding Generalization in Deep Learning

Vyas, Nikhil, Bansal, Yamini, Nakkiran, Preetum

Understand

The ``Neural Tangent Kernel'' (NTK) (Jacot et al 2018), and its empirical variants have been proposed as a proxy to capture certain behaviors of real neural networks.

  • In this work, we study NTKs through the lens of scaling laws, and demonstrate that they fall short of explaining important aspects of neural network generalization.
  • In particular, we demonstrate realistic settings where finite-width neural networks have significantly better data scaling exponents as compared to their corresponding empirical and infinite NTKs at initialization.
  • This reveals a more fundamental difference between the real networks and NTKs, beyond just a few percentage points of test accuracy.

Reading the bibliography…