2018

Deep learning generalizes because the parameter-function map is biased towards simple functions

Valle-Pérez, Guillermo, Camargo, Chico Q., Louis, Ard A.

Understand

Deep neural networks (DNNs) generalize remarkably well without explicit regularization even in the strongly over-parametrized regime where classical learning theory would instead predict that they would severely overfit.

  • While many proposals for some kind of implicit regularization have been made to rationalise this success, there is no consensus for the fundamental reason why DNNs do not strongly overfit.
  • In this paper, we provide a new explanation.
  • By applying a very general probability-complexity bound recently derived from algorithmic information theory (AIT), we argue that the parameter-function map of many DNNs should be exponentially biased towards simple functions.

Reading the bibliography…