Fetching the paper…
Reading the bibliography…
Gradient descent optimization algorithms, while increasingly popular, are often used as black-box optimizers, as practical explanations of their strengths and weaknesses are hard to come by.
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Two problems with backpropagation and other steepest-descent learning procedures for networks, 1986
Richard S. Sutton · 1986
Earlier work this paper cites.
Learning rate schedules for faster stochastic gradient search
C. Darken, J. Chang, and J. Moody · 1992
Earlier work this paper cites.
Efficient BackProp
Yann LeCun, Leon Bottou, Genevieve B. Orr, and Klaus Robert Müller · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Hogwild!: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent
Feng Niu, Benjamin Recht, R Christopher, and Stephen J Wright · 2011
Earlier work this paper cites.
Advances in Optimizing Recurrent Networks
Yoshua Bengio, Nicolas Boulanger-Lewandowski, and Razvan Pascanu · 2012
Cited alongside, same era.
Large Scale Distributed Deep Networks
Jeffrey Dean, Greg S. Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc Aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, and Andrew Y. Ng · 2012
Cited alongside, same era.
ADADELTA: An Adaptive Learning Rate Method
Matthew D. Zeiler · 2012
Cited alongside, same era.
Training Recurrent neural Networks
Ilya Sutskever · 2013
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N. Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Cited alongside, same era.
Delay-Tolerant Algorithms for Asynchronous Distributed Online Learning
Learning to Execute
Wojciech Zaremba and Ilya Sutskever · 2014
Later among the works it cites.
TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems
Martin Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Man, Rajat Monga, Sherry Moore, Derek Murray, Jon Shlens, Benoit Steiner, Ilya Sutskever, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Oriol Vinyals, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Later among the works it cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Sergey Ioffe and Christian Szegedy · 2015
Later among the works it cites.
Adam: a Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Lei Ba · 2015
Later among the works it cites.
Adding Gradient Noise Improves Learning for Very Deep Networks
Arvind Neelakantan, Luke Vilnis, Quoc V. Le, Ilya Sutskever, Lukasz Kaiser, Karol Kurach, and James Martens · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Brendan Mcmahan and Matthew Streeter · 2014
Cited alongside, same era.
Glove: Global Vectors for Word Representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Cited alongside, same era.
A method for unconstrained convex minimization problem with the rate of convergence o(1/k2)
Yurii Nesterov
Cited in the paper.
Later among the works it cites.
Deep learning with Elastic Averaging SGD
Sixin Zhang, Anna Choromanska, and Yann LeCun · 2015
Later among the works it cites.
Incorporating Nesterov Momentum into Adam
Timothy Dozat · 2016
Closest in time.