Fetching the paper…
Reading the bibliography…
In this paper, we theoretically prove that gradient descent can find a global minimum of non-convex optimization of all layers for nonlinear deep neural networks of sizes commonly encountered in practice.
1904
Earlier work this paper cites.
T. M. Cover, “Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition,” IEEE Trans. Electronic Computers , vol. 14, no. 3, pp. 326–334, 1965. [Online]. Available: https://doi.org/10.1109/PGEC.1965.264137
1965
Earlier work this paper cites.
E. B. Baum, “On the capabilities of multilayer perceptrons,” J. Complexity , vol. 4, no. 3, pp. 193–215, 1988. [Online]. Available: https://doi.org/10.1016/0885-064X(88)90020-9
1988
Earlier work this paper cites.
A. Blum and R. L. Rivest, “Training a 3-node neural network is np-complete,” in Advances in neural information processing systems , 1989, pp. 494–501
1989
Earlier work this paper cites.
S. Huang and Y. Huang, “Bounds on the number of hidden neurons in multilayer perceptrons,” IEEE Trans. Neural Networks , vol. 2, no. 1, pp. 47–55, 1991. [Online]. Available: https://doi.org/10.1109/72.80290
1991
Earlier work this paper cites.
M. Yamasaki, “The lower bound of the capacity for a neural network with multiple hidden layers,” in International Conference on Artificial Neural Networks . Springer, 1993, pp. 546–549
1993
Earlier work this paper cites.
G. Huang and H. A. Babri, “Upper bounds on the number of hidden neurons in feedforward networks with arbitrary bounded nonlinear activation functions,” IEEE Trans. Neural Networks , vol. 9, no. 1, pp. 224–229, 1998. [Online]. Available: https://doi.org/10.1109/72.655045
1998
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
P. Bartlett and S. Ben-David, “Hardness results for neural network approximation problems,” in European Conference on Computational Learning Theory . Springer, 1999, pp. 50–62
1999
Cited alongside, same era.
V. Koltchinskii and D. Panchenko, “Empirical margin distributions and bounding the generalization error of combined classifiers,” Annals of Statistics , pp. 1–50, 2002
2002
Cited alongside, same era.
G. Huang, “Learning capability and storage capacity of two-hidden-layer feedforward networks,” IEEE Trans. Neural Networks , vol. 14, no. 2, pp. 274–281, 2003. [Online]. Available: https://doi.org/10.1109/TNN.2003.809401
2003
Cited alongside, same era.
R. Livni, S. Shalev-Shwartz, and O. Shamir, “On the computational efficiency of training neural networks,” in Advances in Neural Information Processing Systems , 2014, pp. 855–863
2014
Cited alongside, same era.
Y. Li and Y. Liang, “Learning overparameterized neural networks via stochastic gradient descent on structured data,” in Advances in Neural Information Processing Systems , 2018, pp. 8157–8166
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
M. Hardt and T. Ma, “Identity matters in deep learning,” in International Conference on Learning Representations , 2017
2017
Cited alongside, same era.
2018
Cited alongside, same era.
Q. Nguyen and M. Hein, “Optimization landscape and expressivity of deep cnns,” in Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018 , 2018, pp. 3727–3736. [Online]. Available: http://proceedings.mlr.press/v80/nguyen18a.html
2018
Cited alongside, same era.
2018
Later among the works it cites.
2019
Closest in time.
2019
Closest in time.