Fetching the paper…
Reading the bibliography…
Although artificial neural networks have shown great promise in applications including computer vision and speech recognition, there remains considerable practical and theoretical difficulty in optimizing their parameters.
W. I. Zangwill, Nonlinear programming : a unified approach, Prentice-Hall international series in management, Prentice-Hall, Englewood Cliffs, N.J., 1969
1969
Earlier work this paper cites.
R. T. Rockafellar, Convex Analysis, Princeton University Press, 1970
1970
Earlier work this paper cites.
doi:10.1287/opre.24.4.643
R. E. Wendell, A. P. Hurter, Minimization of a non-separable objective function subject to disjoint constraints , Operations Research 24 (4) (1976) 643–657 · 1976
Earlier work this paper cites.
R. Johnsonbaugh, W. E. Pfaffenberger, Foundations of Mathematical Analysis, Marcel Dekker, New York, New York, USA, 1981
1981
Earlier work this paper cites.
doi:10.1016/S0893-6080(05)80010-3
A. L. Blum, R. L. Rivest, Training a 3-node neural network is np-complete, Neural Networks 5 (1) (1992) 117–127 · 1992
Earlier work this paper cites.
doi:10.1023/A:1017979506314
I. Tsevendorj, Piecewise-convex maximization problems, Journal of Global Optimization 21 (1) (2001) 1–14 · 2001
Earlier work this paper cites.
S. Ovchinnikov, Max-min representation of piecewise linear functions, Contributions to Algebra and Geometry 43 (1) (2002) 297–302
2002
Cited alongside, same era.
A. N. Iusem, On the convergence properties of the projected gradient method for convex optimization, Computational and Applied Mathematics 22 (2003) 37 – 52
2003
Cited alongside, same era.
S. Boyd, L. Vandenberghe, Convex Optimization, Cambridge University Press, The Edinburgh Building, Cambridge, CB2 8RU, UK, 2004
2004
Cited alongside, same era.
Y. Nesterov, Introductory Lectures on Convex Optimization : A Basic Course, Applied optimization, Kluwer Academic Publishers, Boston, Dordrecht, London, 2004
2004
Cited alongside, same era.
doi:10.1007/s00186-007-0161-1
J. Gorski, F. Pfeuffer, K. Klamroth, Biconvex sets and optimization with biconvex functions: a survey and extensions, Mathematical Methods of Operations Research 66 (3) (2007) 373–407 · 2007
Cited alongside, same era.
2010
Later among the works it cites.
X. Glorot, A. Bordes, Y. Bengio, Deep sparse rectifier neural networks, in: G. J. Gordon, D. B. Dunson (Eds.), Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics (AISTATS-11), Vol. 15, Journal of Machine Learning Research - Workshop and Conference Proceedings, 2011, pp. 315–323
2011
Later among the works it cites.
F. R. Bach, E. Moulines, Non-asymptotic analysis of stochastic approximation algorithms for machine learning, in: Proceedings of the 25th Annual Conference on Neural Information Processing Systems, 2011, pp. 451–459
2011
Later among the works it cites.
doi:10.1016/j.neunet.2012.04.011
P. Baldi, Z. Lu, Complex-valued autoencoders , Neural Networks 33 (2012) 136–147 · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Kawaguchi, Deep learning without poor local minima, arXiv 1605.07110 arXiv:1605.07110
Cited in the paper.
Cited in the paper.
R. Ge, F. Huang, C. Jin, Y. Yuan, Escaping from saddle points - online stochastic gradient for tensor decomposition, Vol. 1, Journal of Machine Learning Research - Workshop and Conference Proceedings, 2015, pp. 1–46
2015
Later among the works it cites.