Fetching the paper…
Reading the bibliography…
The loss surface of deep neural networks has recently attracted interest in the optimization and machine learning communities as a prime example of high-dimensional non-convex problem.
Compressed sensing
Donoho, David L. 2006 · 2006
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Vershynin, Roman. 2010 · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, & Singer, Yoram. 2011 · 2011
Earlier work this paper cites.
Recovery of sparse translation-invariant signals with continuous basis pursuit
Ekanadham, Chaitanya, Tranchina, Daniel, & Simoncelli, Eero P. 2011 · 2011
Earlier work this paper cites.
Lecture 6a Overview of mini–batch gradient descent
Hinton, Geoffrey, Srivastava, N, & Swersky, Kevin. 2012 · 2012
Earlier work this paper cites.
Convex relaxations of structured matrix factorizations
Bach, Francis. 2013 · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, Andrew M, McClelland, James L, & Ganguli, Surya. 2013 · 2013
Earlier work this paper cites.
Compressed sensing off the grid
Tang, Gongguo, Bhaskar, Badri Narayan, Shah, Parikshit, & Recht, Benjamin. 2013 · 2013
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Yann N, Pascanu, Razvan, Gulcehre, Caglar, Cho, Kyunghyun, Ganguli, Surya, & Bengio, Yoshua. 2014 · 2014
Cited alongside, same era.
Qualitatively characterizing neural network optimization problems
Goodfellow, Ian J, Vinyals, Oriol, & Saxe, Andrew M. 2014 · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik, & Ba, Jimmy. 2014 · 2014
Cited alongside, same era.
Explorations on high dimensional landscapes
Sagun, Levent, Guney, V Ugur, Arous, Gerard Ben, & LeCun, Yann. 2014 · 2014
On the quality of the initial basin in overspecified neural networks
Safran, Itay, & Shamir, Ohad. 2015 · 2015
Later among the works it cites.
Deep Learning without Poor Local Minima
Kawaguchi, Kenji. 2016 · 2016
Closest in time.
Gradient descent converges to minimizers
Lee, Jason D, Simchowitz, Max, Jordan, Michael I, & Recht, Benjamin. 2016 · 2016
Closest in time.
Distribution-Specific Hardness of Learning Neural Networks
Shamir, Ohad. 2016 · 2016
Closest in time.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Soudry, Daniel, & Carmon, Yair. 2016 · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The Loss Surfaces of Multilayer Networks
Choromanska, Anna, Henaff, Mikael, Mathieu, Michael, Arous, Gérard Ben, & LeCun, Yann. 2015 · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, Sergey, & Szegedy, Christian. 2015 · 2015
Cited alongside, same era.
Local minima in training of neural networks
Swirszcz, Grzegorz, Czarnecki, Wojciech Marian, & Pascanu, Razvan. 2016 · 2016
Closest in time.
Symmetry-breaking convergence analysis of certain two-layered neural networks with ReLU nonlinearity
Tian, Yuandong. 2017 · 2017
Closest in time.