Fetching the paper…
Reading the bibliography…
While the optimization problem behind deep neural networks is highly non-convex, it is frequently observed in practice that training deep networks seems possible without getting stuck in suboptimal points.
Lectures on H-Cobordism Theorem
Milnor, J · 1965
Earlier work this paper cites.
Elementary classical analysis
Marsden, J. E · 1974
Earlier work this paper cites.
Neural networks and principle component analysis: Learning from examples without local minima
Baldi, P. and Hornik, K · 1988
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
LeCun, Y., Boser, B., Denker, J.S., Henderson, D., Howard, R.E., Hubbard, W., and Jackel, L.D · 1990
Earlier work this paper cites.
On the problem of local minima in backpropagation
Gori, M. and Tesi, A · 1992
Earlier work this paper cites.
On the local minima free condition of backpropagation learning
Yu, X. and Chen, G · 1995
Earlier work this paper cites.
Exponentially many local minima for single neurons
Auer, P., Herbster, M., and Warmuth, M. K · 1996
Earlier work this paper cites.
Successes and failures of backpropagation: A theoretical investigation
Frasconi, P., Gori, M., and Tesi, A · 1997
Earlier work this paper cites.
Overfitting in neural nets: Backpropagation, conjugate gradient, and early stopping
Caruana, R., Lawrence, S., and Giles, L · 2001
Earlier work this paper cites.
A Primer of Real Analytic Functions
Krantz, S. G. and Parks, H. R · 2002
Earlier work this paper cites.
Training a single sigmoidal neuron is hard
Sima, J · 2002
Earlier work this paper cites.
Multiple view geometry in computer vision
Hartley, R. and Zisserman, A · 2004
Earlier work this paper cites.
Deep, big, simple neural nets for handwritten digit recognition
Ciresan, D. C., Meier, U., Gambardella, L. M., and Schmidhuber, J · 2010
Earlier work this paper cites.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
Vincent, P., Larochelle, H., Lajoie, I., Bengio, Y., and Manzagol, P · 2010
Cited alongside, same era.
Bayesian Reasoning and Machine Learning
Barber, D · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Cited alongside, same era.
Do deep nets really need to be deep?
Ba, J. and Caruana, R · 2014
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Cited alongside, same era.
Qualitatively characterizing neural network optimization problems
Goodfellow, I. J., Vinyals, O., and Saxe, A. M · 2015
Deep learning without poor local minima
Kawaguchi, K · 2016
Later among the works it cites.
Gradient descent only converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B · 2016
Later among the works it cites.
How far can we go without convolution: Improving fully-connected networks
Lin, Z., Memisevic, R., and Konda, K · 2016
Later among the works it cites.
On the quality of the initial basin in overspecified networks
Safran, I. and Shamir, O · 2016
Later among the works it cites.
Singularity of the hessian in deep learning, 2016
Sagun, L., Bottou, L., and LeCun, Y · 2016
Later among the works it cites.
Globally optimal gradient descent for a convnet with gaussian inputs, 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Global optimality in tensor factorization, deep learning, and beyond, 2015
Haeffele, B. D. and Vidal, R · 2015
Cited alongside, same era.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Cited alongside, same era.
The zero set of a real analytic function, 2015
Mityagin, B · 2015
Cited alongside, same era.
Path-sgd: Path-normalized optimization in deep neural networks
Neyshabur, B., Salakhutdinov, R. R., and Srebro, N · 2015
Cited alongside, same era.
Complex powers of analytic functions and meromorphic renormalization in qft, 2015
Nguyen, V. D · 2015
Cited alongside, same era.
Globally optimal training of generalized polynomial neural networks with nonlinear spectral methods
Gautier, A., Nguyen, Q., and Hein, M · 2016
Cited alongside, same era.
Brutzkus, A. and Globerson, A · 2017
Closest in time.
Theory ii: Landscape of the empirical risk in deep learning, 2017
Poggio, T. and Liao, Q · 2017
Closest in time.
Piecewise convexity of artificial neural networks, 2017
Rister, B. and Rubin, D. L · 2017
Closest in time.
Learning relus via gradient descent, 2017
Soltanolkotabi, M · 2017
Closest in time.
Exponentially vanishing sub-optimal local minima in multilayer neural networks, 2017
Soudry, D. and Hoffer, E · 2017
Closest in time.
Understanding deep learning requires re-thinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, Oriol · 2017
Closest in time.
The landscape of deep learning algorithms, 2017
Zhou, P. and Feng, J · 2017
Closest in time.