Fetching the paper…
Reading the bibliography…
The past decade has witnessed a successful application of deep learning to solving many challenging problems in machine learning and artificial intelligence.
Ensembles semi-analytiques
Łojasiewicz, S. (1965) · 1965
Earlier work this paper cites.
Topics in Matrix Analysis
Horn, R. (1986) · 1986
Earlier work this paper cites.
An introduction to computing with neural nets
Lippmann, R. (1988) · 1988
Earlier work this paper cites.
Linear learning: Landscapes and algorithms
Baldi, P. (1989) · 1989
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Baldi, P. and Hornik, K. (1989) · 1989
Earlier work this paper cites.
On the problem of local minima in backpropagation
Gori, M. and Tesi, A. (1992) · 1992
Earlier work this paper cites.
On the local minima free condition of backpropagation learning
Yu, X. H. and Chen, G. A. (1995) · 1995
Earlier work this paper cites.
Phase retrieval via wirtinger flow: Theory and algorithms
Candès, E., Li, X., and Soltanolkotabi, M. (2015) · 2007
Earlier work this paper cites.
Complex-valued autoencoders
Baldi, P. and Lu, Z. (2012) · 2012
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Nesterov, Y. (2014) · 2014
Earlier work this paper cites.
Solving random quadratic systems of equations is nearly as easy as solving linear systems
Chen, Y. and Candès, E. (2015) · 2015
Earlier work this paper cites.
Fast low-rank estimation by projected gradient descent: General statistical and algorithmic guarantees
Chen, Y. and Wainwright, M. J. (2015) · 2015
Cited alongside, same era.
The local convexity of solving quadratic equations
White, C. D., Ward, R., and Sanghavi, S. (2015) · 2015
Cited alongside, same era.
A convergent gradient descent algorithm for rank minimization and semidefinite programming from random linear measurements
Zheng, Q. and Lafferty, J. (2015) · 2015
Cited alongside, same era.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A. (2016) · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the Polyak-Łojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M. (2016) · 2016
Convergence analysis for rectangular matrix completion using burer-monteiro factorization and gradient descent
Zheng, Q. and Lafferty, J. (2016) · 2016
Later among the works it cites.
Geometrical properties and accelerated gradient solvers of non-convex phase retrieval
Zhou, Y., Zhang, H., and Liang, Y. (2016) · 2016
Later among the works it cites.
Porcupine neural networks: (almost) all local optima are global
Feizi, S., Javadi, H., Zhang, J., and Tse, D. (2017) · 2017
Closest in time.
Topology and geometry of half-rectified network optimization
Freeman, C. D. and Bruna, J. (2017) · 2017
Closest in time.
Identity matters in deep learning
Hardt, M. and Ma, T. (2017) · 2017
Closest in time.
Depth creates no bad local minima
Lu, H. T. and Kawaguchi, K. (2017) · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep learning without poor local minima
Kawaguchi, K. (2016) · 2016
Cited alongside, same era.
Stochastic variance reduction for nonconvex optimization
Reddi, S., Hefny, A., Sra, S., Póczós, B., and Smola, A. (2016) · 2016
Cited alongside, same era.
Low-rank solutions of linear matrix equations via procrustes flow
Tu, S., Boczar, R., Soltanolkotabi, M., and Recht, B. (2016) · 2016
Cited alongside, same era.
Solving random systems of quadratic equations via truncated generalized gradient flow
Wang, G. and Giannakis, G. (2016) · 2016
Cited alongside, same era.
Diversity leads to generalization in neural networks
Xie, B., Liang, Y., and Song, L. (2016) · 2016
Cited alongside, same era.
Reshaped wirtinger flow for solving quadratic system of equations
Zhang, H. and Liang, Y. (2016) · 2016
Cited alongside, same era.
The loss surface of deep and wide neural networks
Nguyen, Q. and Hein, M. (2017) · 2017
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Soltanolkotabi, M., Javanmard, A., and Lee, J. D. (2017) · 2017
Closest in time.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
Soudry, D. and Hoffer, E. (2017) · 2017
Closest in time.
How regularization affects the critical points in linear networks
Taghvaei, A., Kim, J. W., and Mehta, P. (2017) · 2017
Closest in time.
Global optimality conditions for deep neural networks
Yun, C., Sra, S., and Jadbabaie, A. (2017) · 2017
Closest in time.
Recovery guarantees for one-hidden-layer neural networks
Zhong, K., Song, Z., Jain, P., Bartlett, P. L., and Dhillon, I. S. (2017) · 2017
Closest in time.