Fetching the paper…
Reading the bibliography…
We present the remote stochastic gradient (RSG) method, which computes the gradients at configurable remote observation points, in order to improve the convergence rate and suppress gradient noise at the same time for different curvatures.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
Updating quasi-Newton matrices with limited storage
J. Nocedal · 1980
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L. Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Gradient-based learning applied to document recognition
A. Krizhevsky · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
L. Bottou · 2010
Earlier work this paper cites.
Deep learning via Hessian-free optimization
J. Martens · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method
Matthew D. Zeiler · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Cited alongside, same era.
A stochastic quasi-Newton method for large-scale optimization
R. H Byrd, S. L Hansen, J. Nocedal, and Y. Singer · 2016
Cited alongside, same era.
Incorporating Nesterov momentum into Adam
Timothy Dozat · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
Understanding the role of momentum in stochastic gradient methods
Igor Gitman, Hunter Lang, Pengchuan Zhang, and Lin Xiao · 2019
Closest in time.
Small steps and giant leaps: Minimal newton solvers for deep learning
Joao F. Henriques, Sebastien Ehrhardt, Samuel Albanie, and Andrea Vedaldi · 2019
Closest in time.
Adaptive gradient methods with dynamic bound of learning rate
Liangchen Luo, Yuanhao Xiong, and Yan Liu · 2019
Closest in time.
Quasi-hyperbolic momentum and adam for deep learning
Jerry Ma and Denis Yarats · 2019
Closest in time.
Large-scale distributed second-order optimization using kronecker-factored approximate curvature for deep convolutional neural networks
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka · 2019
Closest in time.
Pytorch: An imperative style, high-performance deep learning library
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nitish Shirish Keskar and Richard Socher · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Cited alongside, same era.
Optimization methods for large-scale machine learning
L. Bottou, F. Curtis, and J. Nocedal · 2018
Cited alongside, same era.
Accelerating stochastic gradient descent for least squares regression
Prateek Jain, Sham M. Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford · 2018
Cited alongside, same era.
On the insufficiency of existing momentum schemes for stochastic optimization
Rahul Kidambi, Praneeth Netrapalli, Prateek Jain, and Sham M. Kakade · 2018
Cited alongside, same era.
On the convergence of Adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
On the convergence of a class of adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2019
Cited alongside, same era.
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Closest in time.
Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM
Qianqian Tong, Guannan Liang, and Jinbo Bi · 2019
Closest in time.
Sadam: A variant of adam for strongly convex functions
Guanghui Wang, Shiyin Lu, Weiwei Tu, and Lijun Zhang · 2019
Closest in time.
Lookahead optimizer: k steps forward, 1 step back
Michael Zhang, James Lucas, Jimmy Ba, and Geoffrey E Hinton · 2019
Closest in time.
Second order optimization made practical
Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer · 2020
Closest in time.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2020
Closest in time.
Gradient descent with momentum — to accelerate or to super-accelerate?
Goran Nakerst, John Brennan, and Masudul Haque · 2020
Closest in time.