Fetching the paper…
Reading the bibliography…
We examine the squared error loss landscape of shallow linear neural networks.
PhD thesis, Harvard University (1974)
Werbos, P.: Beyond regression: new fools for prediction and analysis in the behavioral sciences · 1974
Earlier work this paper cites.
Neural Networks 2
Baldi, P., Hornik, K.: Neural networks and principal component analysis: Learning from examples without local minima · 1989
Earlier work this paper cites.
In: Advances in Neural Information Processing Systems (NIPS), pp. 494–501 (1989)
Blum, A., Rivest, R.L.: Training a 3-node neural network is np-complete · 1989
Earlier work this paper cites.
Mathematics of Control, Signals and Systems 2
Cybenko, G.: Approximation by superpositions of a sigmoidal function · 1989
Earlier work this paper cites.
Neural Networks 2
Hornik, K., Stinchcombe, M., White, H.: Multilayer feedforward networks are universal approximators · 1989
Earlier work this paper cites.
McGraw-Hill New York (1997)
Schalkoff, R.J.: Artificial Neural Networks, vol. 1 · 1997
Earlier work this paper cites.
Mathematical Programming 89
Byrd, R.H., Gilbert, J.C., Nocedal, J.: A trust region method based on interior point techniques for nonlinear programming · 2000
Earlier work this paper cites.
SIAM (2000)
Conn, A.R., Gould, N.I., Toint, P.L.: Trust Region Methods · 2000
Earlier work this paper cites.
Mathematical Programming 108
Nesterov, Y., Polyak, B.T.: Cubic regularization of newton method and its global performance · 2006
Earlier work this paper cites.
Cambridge University Press (2012)
Horn, R.A., Johnson, C.R.: Matrix Analysis · 2012
Earlier work this paper cites.
In: Conference on Learning Theory, pp. 797–842 (2015)
Ge, R., Huang, F., Jin, C., Yuan, Y.: Escaping from saddle points—online stochastic gradient for tensor decomposition · 2015
Earlier work this paper cites.
Nature 521
LeCun, Y., Bengio, Y., Hinton, G.: Deep learning · 2015
Earlier work this paper cites.
In: 53rd Annual Allerton Conference on Communication, Control, and Computing (Allerton), pp. 1336–1343 (2015)
Mousavi, A., Patel, A.B., Baraniuk, R.G.: A deep learning approach to structured signal recovery · 2015
Earlier work this paper cites.
In: Advances in Neural Information Processing Systems, pp. 3873–3881 (2016)
Bhojanapalli, S., Neyshabur, B., Srebro, N.: Global optimality of local search for low rank matrix recovery · 2016
Earlier work this paper cites.
arXiv preprint arXiv:1611.00756 (2016)
Carmon, Y., Duchi, J.C., Hinder, O., Sidford, A.: Accelerated methods for non-convex optimization · 2016
Earlier work this paper cites.
In: Advances in Neural Information Processing Systems, pp. 2973–2981 (2016)
Ge, R., Lee, J.D., Ma, T.: Matrix completion has no spurious local minimum · 2016
Cited alongside, same era.
IEEE Signal Processing Letters 23
Kamilov, U.S., Mansour, H.: Learning optimal nonlinearities for itjin2017escaperative thresholding algorithms · 2016
Cited alongside, same era.
In: Advances in Neural Information Processing Systems, pp. 586–594 (2016)
Kawaguchi, K.: Deep learning without poor local minima · 2016
Cited alongside, same era.
In: Conference on Learning Theory, pp. 1246–1257 (2016)
Lee, J.D., Simchowitz, M., Jordan, M.I., Recht, B.: Gradient descent only converges to minimizers · 2016
Cited alongside, same era.
arXiv preprint arXiv:1612.09296 (2016)
Li, X., Wang, Z., Lu, J., Arora, R., Haupt, J., Liu, H., Zhao, T.: Symmetry, saddle points, and global geometry of nonconvex matrix factorization · 2016
Cited alongside, same era.
arXiv preprint arXiv:1710.07406 (2017)
Lee, J.D., Panageas, I., Piliouras, G., Simchowitz, M., Jordan, M.I., Recht, B.: First-order methods almost always avoid saddle points · 2017
Later among the works it cites.
SIAM Journal on Optimization 27
Liu, H., Yue, M.C., Man-Cho So, A.: On the estimation performance and convergence rate of the generalized power method for phase synchronization · 2017
Later among the works it cites.
arXiv preprint arXiv:1702.08580 (2017)
Lu, H., Kawaguchi, K.: Depth creates no bad local minima · 2017
Later among the works it cites.
In: Artificial Intelligence and Statistics, pp. 65–74 (2017)
Park, D., Kyrillidis, A., Carmanis, C., Sanghavi, S.: Non-square matrix sensing without spurious local minima via the burer-monteiro approach · 2017
Later among the works it cites.
arXiv preprint arXiv:1712.00716 (2017)
Qu, Q., Zhang, Y., Eldar, Y.C., Wright, J.: Convolutional phase retrieval via gradient descent · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
IEEE Geoscience and Remote Sensing Letters 13
Li, Y., Tao, C., Tan, Y., Shang, K., Tian, J.: Unsupervised multilayer feature learning for satellite image scene classification · 2016
Cited alongside, same era.
arXiv preprint arXiv:1605.08361 (2016)
Soudry, D., Carmon, Y.: No bad local minima: Data independent training error guarantees for multilayer neural networks · 2016
Cited alongside, same era.
In: Proceedings of the 49th Annual ACM SIGACT Symposium on Theory of Computing, pp. 1195–1199. ACM (2017)
Agarwal, N., Allen-Zhu, Z., Bullins, B., Hazan, E., Ma, T.: Finding approximate local minima faster than gradient descent · 2017
Cited alongside, same era.
IEEE Transactions on Signal Processing 65
Borgerding, M., Schniter, P., Rangan, S.: Amp-inspired deep networks for sparse linear inverse problems · 2017
Cited alongside, same era.
arXiv preprint arXiv:1703.00412 (2017)
Curtis, F.E., Robinson, D.P.: Exploiting negative curvature in deterministic and stochastic optimization · 2017
Cited alongside, same era.
arXiv preprint arXiv:1709.06129 (2017)
Du, S.S., Lee, J.D., Tian, Y.: When is a convolutional filter easy to learn? · 2017
Cited alongside, same era.
In: International Conference on Machine Learning, pp. 1233–1242 (2017)
Ge, R., Jin, C., Zheng, Y.: No spurious local minima in nonconvex low rank problems: A unified geometric analysis · 2017
Cited alongside, same era.
Later among the works it cites.
arXiv preprint arXiv:1712.08968 (2017)
Safran, I., Shamir, O.: Spurious local minima are common in two-layer relu neural networks · 2017
Later among the works it cites.
arXiv preprint arXiv:1702.05777 (2017)
Soudry, D., Hoffer, E.: Exponentially vanishing sub-optimal local minima in multilayer neural networks · 2017
Later among the works it cites.
IEEE Transactions on Information Theory 63
Sun, J., Qu, Q., Wright, J.: Complete dictionary recovery over the sphere I: Overview and the geometric picture · 2017
Later among the works it cites.
In: International Conference on Machine Learning, pp. 3404–3413 (2017)
Tian, Y.: An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis · 2017
Later among the works it cites.
arXiv preprint arXiv:1707.02444 (2017)
Yun, C., Sra, S., Jadbabaie, A.: Global optimality conditions for deep neural networks · 2017
Later among the works it cites.
arXiv preprint arXiv:1703.01256 (2017)
Zhu, Z., Li, Q., Tang, G., Wakin, M.B.: The global optimization geometry of low-rank matrix optimization · 2017
Later among the works it cites.
to appear in
Li, Q., Zhu, Z., Tang, G.: The non-convex geometry of low-rank matrix optimization · 2018
Closest in time.
IEEE Transactions on Geoscience and Remote Sensing (99), 1–16 (2018)
Li, Y., Zhang, Y., Huang, X., Ma, J.: Learning source-invariant deep hashing convolutional neural networks for cross-source remote sensing image retrieval · 2018
Closest in time.
arXiv preprint arXiv:1803.02968 (2018)
Nouiehed, M., Razaviyayn, M.: Learning deep models: Critical points and local openness · 2018
Closest in time.
IEEE Transactions on Signal Processing 66
Zhu, Z., Li, Q., Tang, G., Wakin, M.B.: Global optimality in low-rank matrix optimization · 2018
Closest in time.