Fetching the paper…
Reading the bibliography…
Deep learning models are often successfully trained using gradient descent, despite the worst case hardness of the underlying non-convex optimization problem.
A polynomial time algorithm that learns two hidden unit nets
Baum, Eric B · 1990
Earlier work this paper cites.
Computers and Intractability; A Guide to the Theory of NP-Completeness
Garey, Michael R. and Johnson, David S · 1990
Earlier work this paper cites.
Training a 3-node neural network is np-complete
Blum, Avrim L and Rivest, Ronald L · 1993
Earlier work this paper cites.
Exponentially many local minima for single neurons
Auer, Peter, Herbster, Mark, Warmuth, Manfred K, et al · 1996
Earlier work this paper cites.
Interval estimation for a binomial proportion
Brown, Lawrence D, Cai, T Tony, and DasGupta, Anirban · 2001
Earlier work this paper cites.
Introductory lectures on convex optimization
Nesterov, Yurii · 2004
Earlier work this paper cites.
Cryptographic hardness of learning
Klivans, Adam · 2008
Earlier work this paper cites.
Kernel methods for deep learning
Cho, Youngmin and Saul, Lawrence K · 2009
Earlier work this paper cites.
Baum’s algorithm learns intersections of halfspaces with respect to log-concave distributions
Klivans, Adam R, Long, Philip M, and Tang, Alex K · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, and Singer, Yoram · 2011
Earlier work this paper cites.
Efficient learning of generalized linear and single index models with isotonic regression
Kakade, Sham M, Kanade, Varun, Shamir, Ohad, and Kalai, Adam · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, Geoffrey, Deng, Li, Yu, Dong, Dahl, George E, Mohamed, Abdel-rahman, Jaitly, Navdeep, Senior, Andrew, Vanhoucke, Vincent, Nguyen, Patrick, Sainath, Tara N, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Cited alongside, same era.
Lin, Min, Chen, Qiang, and Yan, Shuicheng · 2013
Cited alongside, same era.
Learning polynomials with neural networks
Andoni, Alexandr, Panigrahy, Rina, Valiant, Gregory, and Zhang, Li · 2014
Cited alongside, same era.
From average case complexity to improper learning complexity
Daniely, Amit, Linial, Nati, and Shalev-Shwartz, Shai · 2014
Cited alongside, same era.
On the computational efficiency of training neural networks
Livni, Roi, Shalev-Shwartz, Shai, and Shamir, Ohad · 2014
Cited alongside, same era.
The loss surfaces of multilayer networks
Choromanska, Anna, Henaff, Mikael, Mathieu, Michael, Arous, Gérard Ben, and LeCun, Yann · 2015
Gradient descent learns linear dynamical systems
Hardt, Moritz, Ma, Tengyu, and Recht, Benjamin · 2016
Later among the works it cites.
Deep learning without poor local minima
Kawaguchi, Kenji · 2016
Later among the works it cites.
Gradient descent only converges to minimizers
Lee, Jason D., Simchowitz, Max, Jordan, Michael I., and Recht, Benjamin · 2016
Later among the works it cites.
The landscape of empirical risk for non-convex losses
Mei, Song, Bai, Yu, and Montanari, Andrea · 2016
Later among the works it cites.
V-net: Fully convolutional neural networks for volumetric medical image segmentation
Milletari, Fausto, Navab, Nassir, and Ahmadi, Seyed-Ahmad · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Escaping from saddle points-online stochastic gradient for tensor decomposition
Ge, Rong, Huang, Furong, Jin, Chi, and Yuan, Yang · 2015
Cited alongside, same era.
Global optimality in tensor factorization, deep learning, and beyond
Haeffele, Benjamin D and Vidal, René · 2015
Cited alongside, same era.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Janzamin, Majid, Sedghi, Hanie, and Anandkumar, Anima · 2015
Cited alongside, same era.
Variance reduction for faster non-convex optimization
Allen-Zhu, Zeyuan and Hazan, Elad · 2016
Cited alongside, same era.
Matrix completion has no spurious local minimum
Ge, Rong, Lee, Jason D, and Ma, Tengyu · 2016
Cited alongside, same era.
Identity matters in deep learning
Hardt, Moritz and Ma, Tengyu · 2016
Cited alongside, same era.
Safran, Itay and Shamir, Ohad · 2016
Later among the works it cites.
Distribution-specific hardness of learning neural networks
Shamir, Ohad · 2016
Later among the works it cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Soudry, Daniel and Carmon, Yair · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Yonghui, Schuster, Mike, Chen, Zhifeng, Le, Quoc V., Norouzi, Mohammad, Macherey, Wolfgang, Krikun, Maxim, Cao, Yuan, Gao, Qin, Macherey, Klaus, Klingner, Jeff, Shah, Apurva, Johnson, Melvin, Liu, Xiaobing, Kaiser, Lukasz, Gouws, Stephan, Kato, Yoshikiyo, Kudo, Taku, Kazawa, Hideto, Stevens, Keith, Kurian, George, Patil, Nishant, Wang, Wei, Young, Cliff, Smith, Jason, Riesa, Jason, Rudnick, Alex, Vinyals, Oriol, Corrado, Greg, Hughes, Macduff, and Dean, Jeffrey · 2016
Later among the works it cites.
Global analysis of expectation maximization for mixtures of two gaussians
Xu, Ji, Hsu, Daniel J, and Maleki, Arian · 2016
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Zhang, Chiyuan, Bengio, Samy, Hardt, Moritz, Recht, Benjamin, and Vinyals, Oriol · 2016
Later among the works it cites.
Electron-proton dynamics in deep learning
Zhang, Qiuyi, Panigrahy, Rina, Sachdeva, Sushant, and Rahimi, Ali · 2017
Closest in time.