2017

SGD Learns Over-parameterized Networks that Provably Generalize on Linearly Separable Data

Brutzkus, Alon, Globerson, Amir, Malach, Eran et al.

Understand

Neural networks exhibit good generalization behavior in the over-parameterized regime, where the number of network parameters exceeds the number of observations.

  • Nonetheless, current generalization bounds for neural networks fail to explain this phenomenon.
  • In an attempt to bridge this gap, we study the problem of learning a two-layer over-parameterized neural network, when the data is generated by a linearly separable function.
  • In the case where the network has Leaky ReLU activations, we provide both optimization and generalization guarantees for over-parameterized networks.

Reading the bibliography…