Fetching the paper…
Reading the bibliography…
Although neural networks are routinely and successfully trained in practice using simple gradient-based methods, most existing theoretical results are negative, showing that learning such networks is difficult, in a worst-case sense over all data distributions.
Minimization methods for nonsmooth convex and quasiconvex functions
Yurii E Nesterov · 1984
Earlier work this paper cites.
Generalized linear models
Peter McCullagh and John A Nelder · 1989
Earlier work this paper cites.
Weakly learning dnf and characterizing statistical query learning using fourier analysis
Avrim Blum, Merrick Furst, Jeffrey Jackson, Michael Kearns, Yishay Mansour, and Steven Rudich · 1994
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Michael Kearns · 1998
Earlier work this paper cites.
Applied analysis
John K. Hunter and Bruno Nachtergaele · 2001
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Cryptographic hardness for learning intersections of halfspaces
Adam R Klivans and Alexander A Sherstov · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Learning polynomials with neural networks
Alexandr Andoni, Rina Panigrahy, Gregory Valiant, and Li Zhang · 2014
Cited alongside, same era.
Provable bounds for learning some deep representations
Sanjeev Arora, Aditya Bhaskara, Rong Ge, and Tengyu Ma · 2014
Cited alongside, same era.
Embedding hard learning problems into gaussian space
Adam R. Klivans and Pravesh Kothari · 2014
Cited alongside, same era.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Cited alongside, same era.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Cited alongside, same era.
Statistical query algorithms for stochastic convex optimization
Vitaly Feldman, Cristobal Guzman, and Santosh Vempala · 2015
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Later among the works it cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Majid Janzamin, Hanie Sedghi, and Anima Anandkumar · 2015
Later among the works it cites.
Learning halfspaces and neural networks with random initialization
Yuchen Zhang, Jason D Lee, Martin J Wainwright, and Michael I Jordan · 2015
Later among the works it cites.
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Closest in time.
Complexity theoretic limitations on learning dnf’s
Amit Daniely and Shai Shalev-Shwartz · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Beyond convexity: Stochastic quasi-convex optimization
Elad Hazan, Kfir Levy, and Shai Shalev-Shwartz · 2015
Cited alongside, same era.
Closest in time.
On the quality of the initial basin in overspecified neural networks
Itay Safran and Ohad Shamir · 2016
Closest in time.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Closest in time.