Fetching the paper…
Reading the bibliography…
Variance reduction has emerged in recent years as a strong competitor to stochastic gradient descent in non-convex problems, providing the first algorithms to improve upon the converge rate of stochastic gradient descent for finding first-order critical points.
Stochastic approximation method with gradient averaging for unconstrained problems
A. Ruszczynski and W. Syski · 1983
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. C. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Earlier work this paper cites.
Mixed optimization for smooth functions
M. Mahdavi, L. Zhang, and R. Jin · 2013
Earlier work this paper cites.
Variance reduction for stochastic gradient optimization
C. Wang, X. Chen, A. J. Smola, and E. P. Xing · 2013
Earlier work this paper cites.
Linear convergence with condition number independent access of full gradients
L. Zhang, M. Mahdavi, and R. Jin · 2013
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Variance reduction for faster non-convex optimization
Z. Allen-Zhu and E. Hazan · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Stochastic variance reduction for nonconvex optimization
S. J. Reddi, A. Hefny, S. Sra, B. Poczos, and A. Smola · 2016
Cited alongside, same era.
Inexact SARAH algorithm for stochastic optimization
L. M. Nguyen, K. Scheinberg, and M. Takáč · 2018
Later among the works it cites.
On the convergence of Adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Later among the works it cites.
Tensor2tensor for neural machine translation
A. Vaswani, S. Bengio, E. Brevdo, F. Chollet, A. N. Gomez, S. Gouws, L. Jones, Ł. Kaiser, N. Kalchbrenner, N. Parmar, R. Sepassi, N. Shazeer, and J. Uszkoreit · 2018
Later among the works it cites.
AdaGrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization
R. Ward, X. Wu, and L. Bottou · 2018
Later among the works it cites.
Stochastic nested variance reduced gradient descent for nonconvex optimization
D. Zhou, P. Xu, and Q. Gu · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Yuan, B. Ying, and A. H. Sayed · 2016
Cited alongside, same era.
Harder, better, faster, stronger convergence rates for least-squares regression
A. Dieuleveut, N. Flammarion, and F. Bach · 2017
Cited alongside, same era.
Distributed stochastic optimization via adaptive SGD
A. Cutkosky and R. Busa-Fekete · 2018
Cited alongside, same era.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
C. Fang, C. J. Li, Z. Lin, and T. Zhang · 2018
Cited alongside, same era.
Accelerating stochastic gradient descent for least squares regression
P. Jain, S. M. Kakade, R. Kidambi, P. Netrapalli, and A. Sidford · 2018
Cited alongside, same era.
SARAH: A novel method for machine learning problems using stochastic recursive gradient
L. M. Nguyen, J. Liu, K. Scheinberg, and M. Takáč
Cited in the paper.
Stochastic recursive gradient algorithm for nonconvex optimization
L. M. Nguyen, J. Liu, K. Scheinberg, and M. Takáč
Cited in the paper.
Lower bounds for non-convex stochastic optimization
Y. Arjevani, Y. Carmon, J. C. Duchi, D. J. Foster, N. Srebro, and B. Woodworth · 2019
Closest in time.
On the ineffectiveness of variance reduced optimization for deep learning
Aaron Defazio and Léon Bottou · 2019
Closest in time.
On the convergence of stochastic gradient descent with adaptive stepsizes
X. Li and F. Orabona · 2019
Closest in time.
Hybrid stochastic gradient descent algorithms for stochastic nonconvex optimization
Quoc Tran-Dinh, Nhan H Pham, Dzung T Phan, and Lam M Nguyen · 2019
Closest in time.