Fetching the paper…
Reading the bibliography…
We consider the fundamental problem of learning a single neuron $x \mapsto\sigma(w^\top x)$ using standard gradient methods.
Probability inequalities for sums of bounded random variables
W. Hoeffding · 1994
Earlier work this paper cites.
The isotron algorithm: High-dimensional isotonic regression
A. T. Kalai and R. Sastry · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Efficient learning of generalized linear and single index models with isotonic regression
S. M. Kakade, V. Kanade, O. Shamir, and A. Kalai · 2011
Earlier work this paper cites.
A variant of azuma’s inequality for martingales with subgaussian tails
O. Shamir · 2011
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Earlier work this paper cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
M. Janzamin, H. Sedghi, and A. Anandkumar · 2015
Earlier work this paper cites.
A stochastic pca and svd algorithm with an exponential convergence rate
O. Shamir · 2015
Earlier work this paper cites.
When are nonconvex problems not scary?
J. Sun, Q. Qu, and J. Wright · 2015
Earlier work this paper cites.
Reliably learning the relu in polynomial time
S. Goel, V. Kanade, A. Klivans, and J. Thaler · 2016
Cited alongside, same era.
The landscape of empirical risk for non-convex losses
S. Mei, Y. Bai, and A. Montanari · 2016
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
A. Brutzkus and A. Globerson · 2017
Cited alongside, same era.
SGD learns the conjugate kernel class of the network
A. Daniely · 2017
Cited alongside, same era.
When is a convolutional filter easy to learn?
S. S. Du, J. D. Lee, and Y. Tian · 2017
Cited alongside, same era.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
S. Oymak and M. Soltanolkotabi · 2018
Later among the works it cites.
Distribution-specific hardness of learning neural networks
O. Shamir · 2018
Later among the works it cites.
A geometric analysis of phase retrieval
J. Sun, Q. Qu, and J. Wright · 2018
Later among the works it cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Z. Allen-Zhu, Y. Li, and Y. Liang · 2019
Later among the works it cites.
S. Arora, S. S. Du, W. Hu, Z. Li, and R. Wang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
How to escape saddle points efficiently
C. Jin, R. Ge, P. Netrapalli, S. M. Kakade, and M. I. Jordan · 2017
Cited alongside, same era.
Spurious local minima are common in two-layer relu neural networks
I. Safran and O. Shamir · 2017
Cited alongside, same era.
Learning relus via gradient descent
M. Soltanolkotabi · 2017
Cited alongside, same era.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Y. Tian · 2017
Cited alongside, same era.
Y. Cao and Q. Gu · 2019
Later among the works it cites.
Fitting relus via sgd and quantized sgd
S. M. M. Kalan, M. Soltanolkotabi, and A. S. Avestimehr · 2019
Later among the works it cites.
Y. S. Tan and R. Vershynin · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
G. Yehudai and O. Shamir · 2019
Later among the works it cites.