Yoshua Bengio, Nicolas L Roux, Pascal Vincent, Olivier Delalleau, and Patrice Marcotte, Convex neural networks , Advances in neural information processing systems, 2006, pp. 123–130
2006
Earlier work this paper cites.
Sanjeev Arora, Aditya Bhaskara, Rong Ge, and Tengyu Ma, Provable bounds for learning some deep representations , International Conference on Machine Learning, 2014, pp. 584–592
2014
Earlier work this paper cites.
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun, The loss surfaces of multilayer networks , Artificial Intelligence and Statistics, 2015, pp. 192–204
2015
Earlier work this paper cites.
Tamir Hazan and Tommi Jaakkola, Steps toward deep kernel methods from infinite neural networks , arXiv preprint arXiv:1508.05133 (2015)
Original
2015
Earlier work this paper cites.
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton, Deep learning , nature 521
2015
Earlier work this paper cites.
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander Berg, and Fei-Fei Li, Imagenet large scale visual recognition challenge , International Journal of Computer Vision 115
2015
Earlier work this paper cites.
Karen Simonyan and Andrew Zisserman, Very deep convolutional networks for large-scale image recognition , International Conference on Learning Representations (2015)
2015
Earlier work this paper cites.
Hrushikesh N Mhaskar and Tomaso Poggio, Deep vs. shallow networks: An approximation theory perspective , Analysis and Applications 14
2016
Earlier work this paper cites.
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli, Exponential expressivity in deep neural networks through transient chaos , Advances in neural information processing systems, 2016, pp. 3360–3368
2016
Earlier work this paper cites.
Daniel Soudry and Yair Carmon, No bad local minima: Data independent training error guarantees for multilayer neural networks , arXiv preprint arXiv:1605.08361 (2016)
Original
2016
Earlier work this paper cites.
Itay Safran and Ohad Shamir, On the quality of the initial basin in overspecified neural networks , International Conference on Machine Learning, 2016, pp. 774–782
2016
Earlier work this paper cites.
C Daniel Freeman and Joan Bruna, Topology and geometry of half-rectified network optimization , International Conference on Learning Representations (2017)
2017
Earlier work this paper cites.
Quynh Nguyen and Matthias Hein, The loss surface of deep and wide neural networks , Proceedings of the 34th International Conference on Machine Learning, vol. 70, 2017, pp. 2603–2612
2017
Earlier work this paper cites.
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli, Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice , Advances in neural information processing systems, 2017, pp. 4785–4795
2017
Earlier work this paper cites.
Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein, Deep information propagation , International Conference on Learning Representations (2017)
2017
Earlier work this paper cites.
Ravid Shwartz-Ziv and Naftali Tishby, Opening the black box of deep neural networks via information , arXiv preprint arXiv:1703.00810 (2017)
Original
2017
Earlier work this paper cites.
Ge Yang and Samuel Schoenholz, Mean field residual networks: On the edge of chaos , Advances in neural information processing systems, 2017, pp. 7103–7114
2017
Earlier work this paper cites.
Kai Zhong, Zhao Song, Prateek Jain, Peter L. Bartlett, and Inderjit S. Dhillon, Recovery guarantees for one-hidden-layer neural networks , Proceedings of the 34th International Conference on Machine Learning, vol. 70, 2017, pp. 4140–4149
2017
Earlier work this paper cites.