Understand
We explore some mathematical features of the loss landscape of overparameterized neural networks.
- A priori one might imagine that the loss function looks like a typical function from $\mathbb{R}^n$ to $\mathbb{R}$ - in particular, nonconvex, with discrete global minima.
- In this paper, we prove that in at least one important way, the loss function of an overparameterized neural network does not look like a typical function.
- If a neural net has $n$ parameters and is trained on $d$ data points, with $n>d$, we show that the locus $M$ of global minima of $L$ is usually not discrete, but rather an $n-d$ dimensional submanifold of $\mathbb{R}^n$.
Built on
Nothing clear enough to list yet.
Similar
Understanding deep learning requires rethinking generalization
S. Bengio, M. Hardt, B. Recht, O. Vinyals, C. Zhang
Cited in the paper.
Entropy-SGD: Biasing gradient descent into wide valleys
P. Chaudhari, A. Choromanska, S. Soatto,Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, R. Zecchina
Cited in the paper.
Towards Understanding Generalization of Deep Learning: Perspective of Loss Landscapes
L. Wu, Z. Zhu, W. E
Cited in the paper.
Then
Nothing clear enough to list yet.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…