Fetching the paper…
Reading the bibliography…
In deep learning, \textit{depth}, as well as \textit{nonlinearity}, create non-convex loss surfaces.
Perturbation bounds in connection with singular value decomposition
Wedin, P.-Å. (1972) · 1972
Earlier work this paper cites.
Linear learning: Landscapes and algorithms
Baldi, P. (1989) · 1989
Earlier work this paper cites.
Training a 3-node neural network is np-complete
Blum, A. L. and Rivest, R. L. (1992) · 1992
Earlier work this paper cites.
Functionally equivalent feedforward neural networks
Krkova, V. and Kainen, P. C. (1994) · 1994
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Livni, R., Shalev-Shwartz, S., and Shamir, O. (2014) · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S. (2014) · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Ben Arous, G., and LeCun, Y. (2015) · 2015
Cited alongside, same era.
Global optimality in tensor factorization, deep learning, and beyond
Haeffele, B. D. and Vidal, R. (2015) · 2015
Cited alongside, same era.
Bayesian optimization with exponential convergence
Kawaguchi, K., Kaelbling, L. P., and Lozano-Pérez, T. (2015) · 2015
Cited alongside, same era.
Understanding deep neural networks with rectified linear units
Arora, R., Basu, A., Mianjy, P., and Mukherjee, A. (2016) · 2016
Cited alongside, same era.
Topology and geometry of half-rectified network optimization
Freeman, C. D. and Bruna, J. (2016) · 2016
Cited alongside, same era.
Global continuous optimization with error bound and fast convergence
Kawaguchi, K., Maruyama, Y., and Zheng, X. (2016) · 2016
Later among the works it cites.
Learning functions: When is deep better than shallow
Mhaskar, H., Liao, Q., and Poggio, T. (2016) · 2016
Later among the works it cites.
Why and when can deep–but not shallow–networks avoid the curse of dimensionality: a review
Poggio, T., Mhaskar, H., Rosasco, L., Miranda, B., and Liao, Q. (2016) · 2016
Later among the works it cites.
Distribution-specific hardness of learning neural networks
Shamir, O. (2016) · 2016
Later among the works it cites.
Local minima in training of deep networks
Swirszcz, G., Czarnecki, W. M., and Pascanu, R. (2016) · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matrix completion has no spurious local minimum
Ge, R., Lee, J. D., and Ma, T. (2016) · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kawaguchi, K. (2016) · 2016
Cited alongside, same era.
Later among the works it cites.
Optimization as estimation with gaussian processes in bandit settings
Wang, Z., Zhou, B., and Jegelka, S. (2016) · 2016
Later among the works it cites.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
Soudry, D. and Hoffer, E. (2017) · 2017
Closest in time.