2020

Optimization and Generalization of Shallow Neural Networks with Quadratic Activation Functions

Mannelli, Stefano Sarao, Vanden-Eijnden, Eric, Zdeborová, Lenka

Understand

We study the dynamics of optimization and the generalization properties of one-hidden layer neural networks with quadratic activation function in the over-parametrized regime where the layer width $m$ is larger than the input dimension $d$.

  • We consider a teacher-student scenario where the teacher has the same structure as the student with a hidden layer of smaller width $m^*\le m$.
  • We describe how the empirical loss landscape is affected by the number $n$ of data samples and the width $m^*$ of the teacher network.
  • In particular we determine how the probability that there be no spurious minima on the empirical loss depends on $n$, $d$, and $m^*$, thereby establishing conditions under which the neural network can in principle recover the teacher.

Reading the bibliography…