Fetching the paper…
Reading the bibliography…
We study the loss surface of neural networks equipped with a hinge loss criterion and ReLU or leaky ReLU nonlinearities.
On the problem of local minima in backpropagation
Gori, M. and Tesi, A · 1992
Earlier work this paper cites.
Successes and failures of backpropagation: A theoretical
Frasconi, P., Gori, M., and Tesi, A · 1997
Earlier work this paper cites.
Convex analysis and nonlinear optimization: theory and examples. Second Edition
Borwein, J. and Lewis, A. S · 2010
Earlier work this paper cites.
The loss surfaces of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y · 2015
Earlier work this paper cites.
Deep learning without poor local minima
Kawaguchi, K · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
The landscape of empirical risk for non-convex losses
Mei, S., Bai, Y., and Montanari, A · 2016
Cited alongside, same era.
On the quality of the initial basin in overspecified neural networks
Safran, I. and Shamir, O · 2016
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A., Globerson, A., Malach, E., and Shalev-Shwartz, S · 2017
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., and LeCun, Y · 2017
Cited alongside, same era.
When is a convolutional filter easy to learn?
Du, S. S., Lee, J. D., and Tian, Y · 2017
Cited alongside, same era.
Learning relus via gradient descent
Soltanolkotabi, M · 2017
Closest in time.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
Soudry, D. and Hoffer, E · 2017
Closest in time.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Tian, Y · 2017
Closest in time.
Recovery guarantees for one-hidden-layer neural networks
Zhong, K., Song, Z., Jain, P., Bartlett, P. L., and Dhillon, I. S · 2017
Closest in time.
The landscape of deep learning algorithms
Zhou, P. and Feng, J · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convergence analysis of two-layer neural networks with relu activation
Li, Y. and Yuan, Y · 2017
Cited alongside, same era.
Laurent, T. and von Brecht, J. H · 2018
Closest in time.