2022

The Neural Covariance SDE: Shaped Infinite Depth-and-Width Networks at Initialization

Li, Mufan Bill, Nica, Mihai, Roy, Daniel M.

Understand

The logit outputs of a feedforward neural network at initialization are conditionally Gaussian, given a random covariance matrix defined by the penultimate layer.

  • In this work, we study the distribution of this random matrix.
  • Recent work has shown that shaping the activation function as network depth grows large is necessary for this covariance matrix to be non-degenerate.
  • However, the current infinite-width-style understanding of this shaping method is unsatisfactory for large depth: infinite-width analyses ignore the microscopic fluctuations from layer to layer, but these fluctuations accumulate over many layers.

Reading the bibliography…