2017

Learning Non-overlapping Convolutional Neural Networks with Multiple Kernels

Zhong, Kai, Song, Zhao, Dhillon, Inderjit S.

Understand

In this paper, we consider parameter recovery for non-overlapping convolutional neural networks (CNNs) with multiple kernels.

  • We show that when the inputs follow Gaussian distribution and the sample size is sufficiently large, the squared loss of such CNNs is $\mathit{~locally~strongly~convex}$ in a basin of attraction near the global optima for most popular activation functions, like ReLU, Leaky ReLU, Squared ReLU, Sigmoid and Tanh.
  • The required sample complexity is proportional to the dimension of the input and polynomial in the number of kernels and a condition number of the parameters.
  • We also show that tensor methods are able to initialize the parameters to the local strong convex region.

Reading the bibliography…