Fetching the paper…
Reading the bibliography…
In this paper, we study the problem of optimizing a two-layer artificial neural network that best fits a training dataset.
Training a 3-node neural network is np-complete
Blum, A., and Rivest, R. L · 1988
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Barron, A. R · 1994
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Collobert, R., and Weston, J · 2008
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A · 2009
Earlier work this paper cites.
Efficient learning of generalized linear and single index models with isotonic regression
Kakade, S., Kalai, A. T., Kanade, V., and Shamir, O · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
An optimal method for stochastic composite optimization
Lan, G · 2012
Earlier work this paper cites.
Acoustic modeling using deep belief networks
Mohamed, A., Dahl, G. E., and Hinton, G · 2012
Cited alongside, same era.
Stochastic first- and zeroth-order methods for non-convex stochastic programming
Ghadimi, S., and Lan, G · 2013
Cited alongside, same era.
Generalized uniformly optimal methods for nonlinear programming
Ghadimi, S., Lan, G., and Zhang, H · 2015
Cited alongside, same era.
Global optimality in tensor factorization, deep learning, and beyond
Haeffele, B. D., and Vidal, R · 2015
Cited alongside, same era.
Beyond convexity: Stochastic quasi-convex optimization
Hazan, E., Levy, K. Y., and Shalev-Shwartz, S · 2015
Cited alongside, same era.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
Ghadimi, S., and Lan, G · 2016
Cited alongside, same era.
Diversity leads to generalization in neural networks
Xie, B., Liang, Y., and Song, L · 2016
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Later among the works it cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Brutzkus, A., and Globerson, A · 2017
Closest in time.
Convergence analysis of two-layer neural networks with relu activation
Li, Y., and Yuan, Y · 2017
Closest in time.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N · 2017
Closest in time.
The loss surface of deep and wide neural networks
Nguyen, Q. N., and Hein, M · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
No bad local minima: Data independent training error guarantees for multilayer neural networks
Soudry, D., and Carmon, Y · 2016
Cited alongside, same era.
Variance reduction for faster non-convex optimization
Allen-Zhu, Z., and Hazan, E
Cited in the paper.
The loss surface of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y
Cited in the paper.
The isotron algorithm: High-dimensional isotonic regression
Kalai, A., and Sastry, R
Cited in the paper.
The landscape of empirical risk for non-convex losses
Mei, S., Bai, Y., and Montanari, A
Cited in the paper.
Stochastic variance reduction for nonconvex optimization
Reddi, S. J., Hefny, A., Sra, S., Póczós, B., and Smola, A
Cited in the paper.
Theoretical insights into the optimization landscape of over-parametrized shallow neural networks
Soltanolkotabi, M., Javanmard, A., and Lee, J. D · 2017
Closest in time.