Fetching the paper…
Reading the bibliography…
The success of deep neural networks hinges on our ability to accurately and efficiently optimize high-dimensional, non-convex functions.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Coefficients for the study of runge-kutta integration processes
John C. Butcher · 1963
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B.T Polyak · 1964
Earlier work this paper cites.
The convergence of a class of double-rank minimization algorithms 1. general considerations
C. G Broyden · 1970
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o(1/sqr(k))
Yurii Nesterov · 1983
Earlier work this paper cites.
Solving Ordinary Differential Equations I – Nonstiff
Ernst Hairer, Nørsett Syvert P., and Gerhard Wanner · 1987
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and K. Hornik · 1989
Earlier work this paper cites.
Stochastic gradient learning in neural networks
Leon Bottou · 1991
Earlier work this paper cites.
The mnist database of handwritten digits, 1998
Yann LeCun, Corinna Cortes, and Christopher JC Burges · 1998
Earlier work this paper cites.
The statistics of critical points of gaussian fields on large-dimensional spaces
Alan J. Bray and David S. Dean · 2007
Earlier work this paper cites.
Replica symmetry breaking condition exposed by random matrix calculation of landscape complexity
Yan V. Fyodorov and Ian Williams · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Y. Bengio · 2010
Cited alongside, same era.
Deep learning via hessian-free optimization
James Martens · 2010
Cited alongside, same era.
Adaptive deconvolutional networks for mid and high level feature learning
Matthew D. Zeiler, Graham W. Taylor, and Rob Fergus · 2011
Cited alongside, same era.
Rmsprop gradient optimization
Tijmen Tieleman and Geoffery Hinton · 2012
Cited alongside, same era.
On the importance of momentum and initialization in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffery Hinton · 2013
Cited alongside, same era.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaf, Michael Mathieu, Gerard Ben Arous, and Yann LeCun · 2015
Later among the works it cites.
Open problem: The landscape of the loss surfaces of multilayer networks
Anna Choromanska, Yann LeCun, and Gerard Ben Arous · 2015
Later among the works it cites.
Convergence rates of sub-sampled newton methods
Murat A. Erdogdu and Andrea Montanari · 2015
Later among the works it cites.
Qualitatively characterizing neural network optimization problems
Ian J. Goodfellow, Oriol Vinyals, and Andrew M. Saxe · 2015
Later among the works it cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yann N. Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Network in network
Min Lin, Qiang Chen, and Shuicheng Yan · 2014
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, H. Deng, J. adn Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2014
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya. Ganguli · 2014
Cited alongside, same era.
Sergey Ioffe and Christian Szegedy · 2015
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Later among the works it cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Closest in time.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Closest in time.
Probabilistic line searches for stochastic optimization
Giorgio Parisi · 2016
Closest in time.
Local minima in training of deep networks
Grzegorz Swirszcz, Wojciech Marian Czarnecki, and Razvan Pascanu · 2016
Closest in time.