Fetching the paper…
Reading the bibliography…
Momentum plays a crucial role in stochastic gradient-based optimization algorithms for accelerating or improving training deep neural networks (DNNs).
Methods of conjugate gradients for solving linear systems
Magnus R Hestenes et al · 1952
Earlier work this paper cites.
Function minimization by conjugate gradients
Reeves Fletcher and Colin M Reeves · 1964
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris Polyak · 1964
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
Note sur la convergence de méthodes de directions conjuguées
Elijah Polak and Gerard Ribiere · 1969
Earlier work this paper cites.
Nonlinear programming, computational methods, in integer and nonlinear programming
G. Zoutendijk · 1970
Earlier work this paper cites.
Some convergence properties of the conjugate gradient method
Michael James David Powell · 1976
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Yurii E Nesterov · 1983
Earlier work this paper cites.
Descent Property and Global Convergence of the Fletcher—Reeves Method with Inexact Line Search
M. AL-BAALI · 1985
Earlier work this paper cites.
Global convergence properties of conjugate gradient methods for optimization
Jean Charles Gilbert and Jorge Nocedal · 1992
Earlier work this paper cites.
Principles of risk minimization for learning theory
Vladimir Vapnik · 1992
Earlier work this paper cites.
Moller, m.f.: A scaled conjugate gradient algorithm for fast supervised learning. neural networks 6, 525-533
Martin Moller · 1993
Earlier work this paper cites.
An introduction to the conjugate gradient method without the agonizing pain, 1994
Jonathan Richard Shewchuk et al · 1994
Earlier work this paper cites.
Introductory lectures on convex programming volume i: Basic course
Yurii Nesterov · 1998
Earlier work this paper cites.
A nonlinear conjugate gradient method with a strong global convergence property
Yu-Hong Dai and Yaxiang Yuan · 1999
Earlier work this paper cites.
Inexact preconditioned conjugate gradient method with inner-outer iteration
Gene H. Golub and Qiang Ye · 1999
Earlier work this paper cites.
Analysis of the finite precision bi-conjugate gradient algorithm for nonsymmetric linear systems
Charles H. Tong and Qiang Ye · 2000
Cited alongside, same era.
Iterative Methods for Sparse Linear Systems
Yousef Saad · 2003
Cited alongside, same era.
A new conjugate gradient method with guaranteed descent and an efficient line search
William W Hager and Hongchao Zhang · 2005
Cited alongside, same era.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
On optimization methods for deep learning, 2011
Quoc V. Le, Jiquan Ngiam, Adam Coates, Abhik Lahiri, Bobby Prochnow, and Andrew Y. Ng · 2011
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Later among the works it cites.
Towards evaluating the robustness of neural networks
N. Carlini and D.A. Wagner · 2016
Later among the works it cites.
Incorporating nesterov momentum into adam
Timothy Dozat · 2016
Later among the works it cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Cited alongside, same era.
Advances in optimizing recurrent networks
Yoshua Bengio, Nicolas Boulanger-Lewandowski, and Razvan Pascanu · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
First-order methods of smooth convex optimization with inexact oracle
Olivier Devolder, François Glineur, and Yurii Nesterov · 2014
Cited alongside, same era.
Explaining and harnessing adversarial examples
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2014
Cited alongside, same era.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Cited alongside, same era.
Later among the works it cites.
On the generalized lanczos trust-region method
Lei-Hong Zhang, Chungen Shen, and Ren-Cang Li · 2017
Later among the works it cites.
Nonlinear conjugate gradients for scaling synchronous distributed dnn training, 2018
Saurabh Adya, Vinay Palakkode, and Oncel Tuzel · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Later among the works it cites.
Analysis of krylov subspace solutions of regularized non-convex quadratic problems
Yair Carmon and John C Duchi · 2018
Later among the works it cites.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter · 2018
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2019
Later among the works it cites.
ResNet ensemble via the Feynman-Kac formalism to improve natural and robust acurcies
B. Wang, B. Yuan, Z. Shi, and S. Osher · 2019
Later among the works it cites.
Scheduled restart momentum for accelerated stochastic gradient descent
Bao Wang, Tan M Nguyen, Andrea L Bertozzi, Richard G Baraniuk, and Stanley J Osher · 2020
Closest in time.