Fetching the paper…
Reading the bibliography…
We consider the teacher-student setting of learning shallow neural networks with quadratic activations and planted weight matrix $W^*\in\mathbb{R}^{m\times d}$, where $m$ is the width of the hidden layer and $d\le m$ is the data dimension.
Walter Rudin et al., Principles of mathematical analysis , vol. 3, McGraw-hill New York, 1964
1964
Earlier work this paper cites.
Zhidong D Bai and Yong Q Yin, Convergence to the semicircle law , The Annals of Probability (1988), 863–875
1988
Earlier work this paper cites.
Avrim Blum and Ronald L Rivest, Training a 3-node neural network is np-complete , Advances in neural information processing systems, 1989, pp. 494–501
1989
Earlier work this paper cites.
Timothy Poston, C-N Lee, Y Choie, and Yonghoon Kwon, Local minima and back propagation , IJCNN-91-Seattle International Joint Conference on Neural Networks, vol. 2, IEEE, 1991, pp. 173–176
1991
Earlier work this paper cites.
ZD Bai, YQ Yin, et al., Limit of the smallest eigenvalue of a large dimensional sample covariance matrix , The Annals of Probability 21
1993
Earlier work this paper cites.
Andrew R Barron, Approximation and estimation bounds for artificial neural networks , Machine learning 14
1994
Earlier work this paper cites.
Richard Caron and Tim Traynor, The zero set of a polynomial , WSMR Report (2005), 05–02
2005
Earlier work this paper cites.
Guang-Bin Huang, Qin-Yu Zhu, and Chee-Kheong Siew, Extreme learning machine: theory and applications , Neurocomputing 70
2006
Earlier work this paper cites.
Ronan Collobert and Jason Weston, A unified architecture for natural language processing: Deep neural networks with multitask learning , Proceedings of the 25th international conference on Machine learning, ACM, 2008, pp. 160–167
2008
Earlier work this paper cites.
Ali Rahimi and Benjamin Recht, Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning , Advances in neural information processing systems, 2009, pp. 1313–1320
2009
Earlier work this paper cites.
2010
Earlier work this paper cites.
Emmanuel J Candes and Yaniv Plan, Tight oracle inequalities for low-rank matrix recovery from a minimal number of noisy random measurements , IEEE Transactions on Information Theory 57
2011
Earlier work this paper cites.
Abdel-rahman Mohamed, George E Dahl, and Geoffrey Hinton, Acoustic modeling using deep belief networks , IEEE transactions on audio, speech, and language processing 20
2011
Earlier work this paper cites.
Roger A Horn and Charles R Johnson, Matrix analysis , Cambridge University Press, 2012
2012
Earlier work this paper cites.
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton, Imagenet classification with deep convolutional neural networks , Advances in neural information processing systems, 2012, pp. 1097–1105
2012
Earlier work this paper cites.
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir, On the computational efficiency of training neural networks , Advances in neural information processing systems, 2014, pp. 855–863
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun, The loss surfaces of multilayer networks , Artificial Intelligence and Statistics, 2015, pp. 192–204
2015
Earlier work this paper cites.
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan, Escaping from saddle points—online stochastic gradient for tensor decomposition , Conference on Learning Theory, 2015, pp. 797–842
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Benjamin Haeffele, Eric Young, and Rene Vidal, Structured low-rank matrix factorization: Optimality, algorithm, and applications to image processing , International conference on machine learning, 2014, pp. 2007–2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro, Norm-based capacity control in neural networks , Conference on Learning Theory, 2015, pp. 1376–1401
2015
Earlier work this paper cites.
Ronen Eldan and Ohad Shamir, The power of depth for feedforward neural networks , Conference on learning theory, 2016, pp. 907–940
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, Deep residual learning for image recognition , Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778
2016
Earlier work this paper cites.
Kenji Kawaguchi, Deep learning without poor local minima , Advances in neural information processing systems, 2016, pp. 586–594
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht, Gradient descent only converges to minimizers , Conference on learning theory, 2016, pp. 1246–1257
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2017
Later among the works it cites.
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al., Mastering the game of go without human knowledge , Nature 550
2017
Later among the works it cites.
Yuandong Tian, An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis , Proceedings of the 34th International Conference on Machine Learning-Volume 70, JMLR. org, 2017, pp. 3404–3413
2017
Later among the works it cites.
E Weinan, Jiequn Han, and Arnulf Jentzen, Deep learning-based numerical methods for high-dimensional parabolic partial differential equations and backward stochastic differential equations , Communications in Mathematics and Statistics 5
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
Matus Telgarsky, Benefits of depth in neural networks , arXiv preprint arXiv:1602.04485 (2016)
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky, Spectrally-normalized margin bounds for neural networks , Advances in Neural Information Processing Systems, 2017, pp. 6240–6249
2017
Cited alongside, same era.
Alon Brutzkus and Amir Globerson, Globally optimal gradient descent for a convnet with gaussian inputs , Proceedings of the 34th International Conference on Machine Learning-Volume 70, JMLR. org, 2017, pp. 605–614
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Simon S Du, Chi Jin, Jason D Lee, Michael I Jordan, Aarti Singh, and Barnabas Poczos, Gradient descent can take exponential time to escape saddle points , Advances in neural information processing systems, 2017, pp. 1067–1077
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon, Recovery guarantees for one-hidden-layer neural networks , Proceedings of the 34th International Conference on Machine Learning-Volume 70, JMLR. org, 2017, pp. 4140–4149
2017
Later among the works it cites.
2018
Later among the works it cites.
Lenaic Chizat and Francis Bach, On the global convergence of gradient descent for over-parameterized models using optimal transport , Advances in neural information processing systems, 2018, pp. 3036–3046
2018
Later among the works it cites.
Tian Qi Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud, Neural ordinary differential equations , Advances in neural information processing systems, 2018, pp. 6571–6583
2018
Later among the works it cites.
Jeffrey De Fauw, Joseph R Ledsam, Bernardino Romera-Paredes, Stanislav Nikolov, Nenad Tomasev, Sam Blackwell, Harry Askham, Xavier Glorot, Brendan O’Donoghue, Daniel Visentin, et al., Clinically applicable deep learning for diagnosis and referral in retinal disease , Nature medicine 24
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee, Theoretical insights into the optimization landscape of over-parameterized shallow neural networks , IEEE Transactions on Information Theory 65
2018
Later among the works it cites.
Mei Song, Andrea Montanari, and P Nguyen, A mean field view of the landscape of two-layers neural networks , Proceedings of the National Academy of Sciences 115
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2019
Closest in time.
Helmut Bölcskei, Philipp Grohs, Gitta Kutyniok, and Philipp Petersen, Optimal approximation with sparsely connected deep neural networks , SIAM Journal on Mathematics of Data Science 1
2019
Closest in time.
Peter L Bartlett, Nick Harvey, Christopher Liaw, and Abbas Mehrabian, Nearly-tight vc-dimension and pseudodimension bounds for piecewise linear neural networks. , Journal of Machine Learning Research 20
2019
Closest in time.
Justin Sirignano and Konstantinos Spiliopoulos, Mean field analysis of neural networks: A central limit theorem , Stochastic Processes and their Applications (2019)
2019
Closest in time.
2020
Closest in time.