Fetching the paper…
Reading the bibliography…
We study model recovery for data classification, where the training labels are generated from a one-hidden-layer neural network with sigmoid activations, also known as a single-layer feedforward network, and the goal is to recover the weights of the neural network.
G. Cybenko, “Approximation by superpositions of a sigmoidal function,” Mathematics of control, signals and systems , vol. 2, no. 4, pp. 303–314, 1989
1989
Earlier work this paper cites.
K. Hornik, M. Stinchcombe, and H. White, “Multilayer feedforward networks are universal approximators,” Neural networks , vol. 2, no. 5, pp. 359–366, 1989
1989
Earlier work this paper cites.
A. R. Barron, “Universal approximation bounds for superpositions of a sigmoidal function,” IEEE Transactions on Information theory , vol. 39, no. 3, pp. 930–945, 1993
1993
Earlier work this paper cites.
R. Vershynin, “Introduction to the non-asymptotic analysis of random matrices,” Compressed Sensing, Theory and Applications , pp. 210 – 268, 2012
2012
Earlier work this paper cites.
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio, “Identifying and attacking the saddle point problem in high-dimensional non-convex optimization,” in Advances in neural information processing systems , 2014, pp. 2933–2941
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” International Journal of Computer Vision , vol. 115, no. 3, pp. 211–252, 2015
2015
Earlier work this paper cites.
E. J. Candès, X. Li, and M. Soltanolkotabi, “Phase retrieval via Wirtinger flow: Theory and algorithms,” IEEE Transactions on Information Theory , vol. 61, no. 4, pp. 1985–2007, April 2015
2015
Earlier work this paper cites.
J. Sun, Q. Qu, and J. Wright, “Complete dictionary recovery using nonconvex optimization,” International Conference on Machine Learning , pp. 2351–2360, 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al. , “Mastering the game of go with deep neural networks and tree search,” nature , vol. 529, no. 7587, pp. 484–489, 2016
2016
Earlier work this paper cites.
M. Telgarsky, “benefits of depth in neural networks,” in Conference on Learning Theory , 2016, pp. 1517–1539
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
R. Sun and Z.-Q. Luo, “Guaranteed matrix completion via non-convex factorization,” IEEE Transactions on Information Theory , vol. 62, no. 11, pp. 6535–6579, 2016
2016
Earlier work this paper cites.
R. Ge, J. D. Lee, and T. Ma, “Matrix completion has no spurious local minimum,” in Advances in Neural Information Processing Systems , 2016, pp. 2973–2981
2016
Earlier work this paper cites.
S. Bhojanapalli, B. Neyshabur, and N. Srebro, “Global optimality of local search for low rank matrix recovery,” in Advances in Neural Information Processing Systems , 2016, pp. 3873–3881
2016
Cited alongside, same era.
I. Safran and O. Shamir, “On the quality of the initial basin in overspecified neural networks,” in International Conference on Machine Learning , 2016, pp. 774–782
2016
Cited alongside, same era.
P. L. Bartlett, D. J. Foster, and M. J. Telgarsky, “Spectrally-normalized margin bounds for neural networks,” in Advances in Neural Information Processing Systems , 2017, pp. 6241–6250
2017
Cited alongside, same era.
M. Soltanolkotabi, “Learning relus via gradient descent,” in Advances in Neural Information Processing Systems , 2017, pp. 2007–2017
2017
Cited alongside, same era.
S. Oymak, “Learning compact neural networks with regularization,” in Proceedings of the 35th International Conference on Machine Learning , 2018, pp. 3966–3975
2018
Closest in time.
S. S. Du, J. D. Lee, and Y. Tian, “When is a convolutional filter easy to learn?” in International Conference on Learning Representations , 2018
2018
Closest in time.
Y. Chen and Y. Chi, “Harnessing structures in big data via guaranteed low-rank matrix estimation,” IEEE Signal Processing Magazine , 2018
2018
Closest in time.
C. Ma, K. Wang, Y. Chi, and Y. Chen, “Implicit regularization in nonconvex statistical estimation: Gradient descent converges linearly for phase retrieval and matrix completion,” in Proceedings of the 35th International Conference on Machine Learning , 2018, pp. 3345–3354
2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
Y. Li and Y. Yuan, “Convergence analysis of two-layer neural networks with relu activation,” in Advances in Neural Information Processing Systems , 2017, pp. 597–607
2017
Cited alongside, same era.
A. Brutzkus and A. Globerson, “Globally optimal gradient descent for a ConvNet with Gaussian inputs,” in Proceedings of the 34th International Conference on Machine Learning , 2017, pp. 605–614
2017
Cited alongside, same era.
R. Ge and T. Ma, “On the optimization landscape of tensor decompositions,” in Advances in Neural Information Processing Systems 30 , 2017, pp. 3653–3663
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Q. Nguyen and M. Hein, “The loss surface of deep and wide neural networks,” in International Conference on Machine Learning , 2017, pp. 2603–2612
2017
Cited alongside, same era.
Y. Tian, “An analytical formula of population gradient for two-layered ReLU network and its applications in convergence and critical point analysis,” in Proceedings of the 34th International Conference on Machine Learning , 2017, pp. 3404–3413
2017
Cited alongside, same era.
2017
Cited alongside, same era.
M. Soltanolkotabi, A. Javanmard, and J. D. Lee, “Theoretical insights into the optimization landscape of over-parameterized shallow neural networks,” IEEE Transactions on Information Theory , 2018
2018
Closest in time.
R. Ge, J. D. Lee, and T. Ma, “Learning one-hidden-layer neural networks with landscape design,” in International Conference on Learning Representations , 2018
2018
Closest in time.
I. Safran and O. Shamir, “Spurious local minima are common in two-layer ReLU neural networks,” in Proceedings of the 35th International Conference on Machine Learning , 2018, pp. 4433–4441
2018
Closest in time.
2018
Closest in time.
S. Goel, A. Klivans, and R. Meka, “Learning one convolutional layer with overlapping patches,” in Proceedings of the 35th International Conference on Machine Learning , 2018, pp. 1783–1791
2018
Closest in time.
S. Feizi, H. Javadi, J. Zhang, and D. Tse, “Porcupine neural networks: Approximating neural network landscapes,” in Advances in Neural Information Processing Systems 31 , 2018, pp. 4836–4846
2018
Closest in time.
M. Mondelli and A. Montanari, “On the connection between learning two-layer neural networks and tensor decomposition,” in Proceedings of Machine Learning Research , 2019, pp. 1051–1060
2019
Closest in time.
Y. Chi, Y. M. Lu, and Y. Chen, “Nonconvex optimization meets low-rank matrix factorization: An overview,” IEEE Transactions on Signal Processing , vol. 67, no. 20, pp. 5239–5269, 2019
2019
Closest in time.
Y. Chen, Y. Chi, J. Fan, and C. Ma, “Gradient descent with random initialization: Fast global convergence for nonconvex phase retrieval,” Math. Program. 176, 5–37 , 2019
2019
Closest in time.
X. Zhang, Y. Yu, L. Wang, and Q. Gu, “Learning one-hidden-layer relu networks via gradient descent,” International Conference on Artificial Intelligence and Statistics , 2019
2019
Closest in time.