2020

On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime

Wojtowytsch, Stephan

Understand

We describe a necessary and sufficient condition for the convergence to minimum Bayes risk when training two-layer ReLU-networks by gradient descent in the mean field regime with omni-directional initial parameter distribution.

  • This article extends recent results of Chizat and Bach to ReLU-activated networks and to the situation in which there are no parameters which exactly achieve MBR.
  • The condition does not depend on the initalization of parameters and concerns only the weak convergence of the realization of the neural network, not its parameter distribution.

Reading the bibliography…