Fetching the paper…
Reading the bibliography…
Recent years have seen a growing interest in understanding deep neural networks from an optimization perspective.
Wigner, E.P.: On the Distribution of the Roots of Certain Symmetric Matrices. The Annals of Mathematics 67 (1958)
1958
Earlier work this paper cites.
Baldi, P., Hornik, K.: Neural networks and principal component analysis: Learning from examples without local minima. Neural Netw. (1989)
1989
Earlier work this paper cites.
Wright, S., Nocedal, J.: Numerical optimization. Springer Science 35, 67–68 (1999)
1999
Earlier work this paper cites.
Bray, A.J., Dean, D.S.: Statistics of critical points of Gaussian fields on large-dimensional spaces. Physical Review Letters 98, 150201 (2007), https://hal.archives-ouvertes.fr/hal-00124320 , 5 pages
2007
Earlier work this paper cites.
Watanabe, S.: Almost all learning machines are singular. In: 2007 IEEE Symposium on Foundations of Computational Intelligence. pp. 383–388 (2007)
2007
Earlier work this paper cites.
Glorot, X., Bengio, Y.: Understanding the difficulty of training deep feedforward neural networks. In: In Proceedings of the International Conference on Artificial Intelligence and Statistics (AISTATS’10). Society for Artificial Intelligence and Statistics (2010)
2010
Earlier work this paper cites.
Duchi, J., Hazan, E., Singer, Y.: Adaptive subgradient methods for online learning and stochastic optimization. J. Mach. Learn. Res. 12, 2121–2159 (Jul 2011)
2011
Earlier work this paper cites.
2013
Earlier work this paper cites.
Sutskever, I., Martens, J., Dahl, G., Hinton, G.: On the importance of initialization and momentum in deep learning. In: Proceedings of the 30th International Conference on International Conference on Machine Learning - Volume 28. ICML’13 (2013)
2013
Earlier work this paper cites.
Dauphin, Y.N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., Bengio, Y.: Identifying and attacking the saddle point problem in high-dimensional non-convex optimization. In: Proceedings of the 27th International Conference on Neural Information Processing Systems. NIPS’14 (2014)
2014
Cited alongside, same era.
2014
Cited alongside, same era.
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. CoRR abs/1412.6980 (2014)
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2015
Later among the works it cites.
Anandkumar, A., Ge, R.: Efficient approaches for escaping higher order saddle points in non-convex optimization. In: Proceedings of the 29th Conference on Learning Theory, COLT 2016, New York, USA, June 23-26, 2016. pp. 81–102 (2016)
2016
Later among the works it cites.
Goodfellow, I., Bengio, Y., Courville, A.: Deep learning. MIT Press (2016)
2016
Later among the works it cites.
Hardt, M., Recht, B., Singer, Y.: Train faster, generalize better: Stability of stochastic gradient descent. In: Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016. pp. 1225–1234 (2016)
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
2015
Cited alongside, same era.
Choromanska, A., Henaff, M., Mathieu, M., Arous, G.B., LeCun, Y.: The loss surfaces of multilayer networks. In: Proceedings of the Eighteenth International Conference on Artificial Intelligence and Statistics, AISTATS 2015, San Diego, California, USA, May 9-12, 2015 (2015)
2015
Cited alongside, same era.
Ge, R., Huang, F., Jin, C., Yuan, Y.: Escaping from saddle points - online stochastic gradient for tensor decomposition. In: Proceedings of The 28th Conference on Learning Theory, COLT 2015, Paris, France, July 3-6, 2015. pp. 797–842 (2015)
2015
Cited alongside, same era.
Kawaguchi, K.: Deep learning without poor local minima. In: Lee, D.D., Sugiyama, M., Luxburg, U.V., Guyon, I., Garnett, R. (eds.) Advances in Neural Information Processing Systems 29, pp. 586–594. Curran Associates, Inc. (2016)
2016
Later among the works it cites.
Lee, J.D., Simchowitz, M., Jordan, M.I., Recht, B.: Gradient descent only converges to minimizers. In: Proceedings of the 29th Conference on Learning Theory, COLT 2016, New York, USA, June 23-26, 2016. pp. 1246–1257 (2016)
2016
Later among the works it cites.
Zagoruyko, S., Komodakis, N.: Wide residual networks. In: BMVC (2016)
2016
Later among the works it cites.