2024

A Dynamical Model of Neural Scaling Laws

Bordelon, Blake, Atanasov, Alexander, Pehlevan, Cengiz

Understand

On a variety of tasks, the performance of neural networks predictably improves with training time, dataset size and model size across many orders of magnitude.

  • This phenomenon is known as a neural scaling law.
  • Of fundamental importance is the compute-optimal scaling law, which reports the performance as a function of units of compute when choosing model sizes optimally.
  • We analyze a random feature model trained with gradient descent as a solvable model of network training and generalization.

Reading the bibliography…