Fetching the paper…
Reading the bibliography…
Significant advances have been made recently on training neural networks, where the main challenge is in solving an optimization problem with abundant critical points.
Learning polynomials with neural networks
Andoni, A., Panigrahy, R., Valiant, G., and Zhang, L. (2014) · 1916
Earlier work this paper cites.
A bound for the error in the normal approximation to the distribution of a sum of dependent random variables
Stein, C. et al. (1972) · 1972
Earlier work this paper cites.
The estimation of the gradient of a density function, with applications in pattern recognition
Fukunaga, K. and Hostetler, L. (1975) · 1975
Earlier work this paper cites.
Optimal rates of convergence for nonparametric estimators
Stone, C. J. (1980) · 1980
Earlier work this paper cites.
Locally parametric nonparametric density estimation
Hjort, N. L. and Jones, M. (1996) · 1996
Earlier work this paper cites.
Local likelihood density estimation
Loader, C. R. et al. (1996) · 1996
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Hyvärinen, A. (2005) · 2005
Earlier work this paper cites.
On autoencoders and score matching for energy based models
Swersky, K., Buchman, D., Freitas, N. D., Marlin, B. M., et al. (2011) · 2011
Earlier work this paper cites.
Most tensor problems are np-hard
Hillar, C. J. and Lim, L.-H. (2013) · 2013
Earlier work this paper cites.
Provable bounds for learning some deep representations
Arora, S., Bhaskara, A., Ge, R., and Ma, T. (2014) · 2014
Earlier work this paper cites.
Score function features for discriminative learning: Matrix and tensor framework
Janzamin, M., Sedghi, H., and Anandkumar, A. (2014) · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Livni, R., Shalev-Shwartz, S., and Shamir, O. (2014) · 2014
Cited alongside, same era.
Provable tensor methods for learning mixtures of classifiers
Sedghi, H. and Anandkumar, A. (2014) · 2014
Cited alongside, same era.
The loss surfaces of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y. (2015) · 2015
Cited alongside, same era.
Topology and geometry of half-rectified network optimization
Freeman, C. D. and Bruna, J. (2016) · 2016
Cited alongside, same era.
Breaking the bandwidth barrier: Geometrical adaptive entropy estimation
Gao, W., Oh, S., and Viswanath, P. (2016) · 2016
Cited alongside, same era.
Theoretical properties of the global optimizer of two layer neural network
The loss surface and expressivity of deep convolutional neural networks
Nguyen, Q. and Hein, M. (2017) · 2017
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Soltanolkotabi, M., Javanmard, A., and Lee, J. D. (2017) · 2017
Later among the works it cites.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
Soudry, D. and Hoffer, E. (2017) · 2017
Later among the works it cites.
Tian, Y. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Boob, D. and Lan, G. (2017) · 2017
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
Brutzkus, A. and Globerson, A. (2017) · 2017
Cited alongside, same era.
Learning one-hidden-layer neural networks with landscape design
Ge, R., Lee, J. D., and Ma, T. (2017) · 2017
Cited alongside, same era.
Learning depth-three neural networks in polynomial time
Goel, S. and Klivans, A. (2017) · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with relu activation
Li, Y. and Yuan, Y. (2017) · 2017
Cited alongside, same era.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Janzamin, M., Sedghi, H., and Anandkumar, A. (2015a)
Cited in the paper.
Feast at play: Feature extraction using score function tensors
Janzamin, M., Sedghi, H., Niranjan, U., and Anandkumar, A. (2015b)
Cited in the paper.
Zhong, K., Song, Z., Jain, P., Bartlett, P. L., and Dhillon, I. S. (2017) · 2017
Later among the works it cites.
On the power of over-parametrization in neural networks with quadratic activation
Du, S. S. and Lee, J. D. (2018) · 2018
Closest in time.
Learning two-layer neural networks with symmetric inputs
Ge, R., Kuditipudi, R., Li, Z., and Wang, X. (2018) · 2018
Closest in time.
Understanding the loss surface of neural networks for binary classification
Liang, S., Sun, R., Li, Y., and Srikant, R. (2018) · 2018
Closest in time.
Convergence results for neural networks via electrodynamics
Panigrahy, R., Rahimi, A., Sachdeva, S., and Zhang, Q. (2018) · 2018
Closest in time.
Learning one-hidden-layer relu networks via gradient descent
Zhang, X., Yu, Y., Wang, L., and Gu, Q. (2018) · 2018
Closest in time.