Fetching the paper…
Reading the bibliography…
In this paper, we propose a novel stochastic gradient estimator -- ProbAbilistic Gradient Estimator (PAGE) -- for nonconvex optimization.
Gradient methods for the minimisation of functionals
Polyak, B. T · 1963
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence o ( 1 / k 2 ) o(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Nesterov, Y · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A · 2009
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S · 2014
Earlier work this paper cites.
Understanding machine learning: from theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
An optimal randomized incremental gradient method
Lan, G. and Zhou, Y · 2015
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Earlier work this paper cites.
A universal catalyst for first-order optimization
Lin, H., Mairal, J., and Harchaoui, Z · 2015
Earlier work this paper cites.
Variance reduction for faster non-convex optimization
Allen-Zhu, Z. and Hazan, E · 2016
Earlier work this paper cites.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Ghadimi, S., Lan, G., and Zhang, H · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Stochastic variance reduction for nonconvex optimization
Reddi, S. J., Hefny, A., Sra, S., Póczos, B., and Smola, A · 2016
Cited alongside, same era.
Tight complexity bounds for optimizing composite objectives
Woodworth, B. E. and Srebro, N · 2016
Cited alongside, same era.
Katyusha: the first direct acceleration of stochastic gradient methods
Allen-Zhu, Z · 2017
Cited alongside, same era.
Non-convex optimization for machine learning
Jain, P. and Kar, P · 2017
Cited alongside, same era.
Non-convex finite-sum optimization via SCSG methods
Lei, L., Ju, C., Chen, J., and Jordan, M. I · 2017
Cited alongside, same era.
SARAH: A novel method for machine learning problems using stochastic recursive gradient
Nguyen, L. M., Liu, J., Scheinberg, K., and Takáč, M · 2017
Cited alongside, same era.
A unified variance-reduced accelerated gradient method for convex optimization
Lan, G., Li, Z., and Zhou, Y · 2019
Later among the works it cites.
SSRGD: Simple stochastic recursive gradient descent for escaping saddle points
Li, Z · 2019
Later among the works it cites.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Later among the works it cites.
ProxSARAH: An efficient algorithmic framework for stochastic composite nonconvex optimization
Pham, N. H., Nguyen, L. M., Phan, D. T., and Tran-Dinh, Q · 2019
Later among the works it cites.
L-SVRG and L-Katyusha with arbitrary sampling
Qian, X., Qu, Z., and Richtárik, P · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Natasha 2: Faster non-convex optimization than SGD
Allen-Zhu, Z · 2018
Cited alongside, same era.
Improving SAGA via a probabilistic interpolation with gradient descent
Bibi, A., Sailanbayev, A., Ghanem, B., Gower, R. M., and Richtárik, P · 2018
Cited alongside, same era.
SPIDER: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Fang, C., Li, C. J., Lin, Z., and Zhang, T · 2018
Cited alongside, same era.
Random gradient extrapolation for distributed and stochastic optimization
Lan, G. and Zhou, Y · 2018
Cited alongside, same era.
A simple proximal stochastic gradient method for nonsmooth nonconvex optimization
Li, Z. and Li, J · 2018
Cited alongside, same era.
Spiderboost: A class of faster variance-reduced algorithms for nonconvex optimization
Wang, Z., Ji, K., Zhou, Y., Liang, Y., and Tarokh, V · 2018
Cited alongside, same era.
A general analysis framework of lower complexity bounds for finite-sum optimization
Xie, G., Luo, L., and Zhang, Z · 2019
Later among the works it cites.
Lower bounds for smooth nonconvex finite-sum optimization
Zhou, D. and Gu, Q · 2019
Later among the works it cites.
Adaptivity of stochastic gradient methods for nonconvex optimization
Horváth, S., Lei, L., Richtárik, P., and Jordan, M. I · 2020
Closest in time.
Better theory for SGD in the nonconvex world
Khaled, A. and Richtárik, P · 2020
Closest in time.
Don’t jump through hoops and remove those loops: SVRG and Katyusha are better without the outer loop
Kovalev, D., Horváth, S., and Richtárik, P · 2020
Closest in time.
A fast Anderson-Chebyshev acceleration for nonlinear optimization
Li, Z. and Li, J · 2020
Closest in time.
A unified analysis of stochastic gradient methods for nonconvex federated optimization
Li, Z. and Richtárik, P · 2020
Closest in time.
MARINA: Faster non-convex distributed learning with compression
Gorbunov, E., Burlachenko, K., Li, Z., and Richtárik, P · 2021
Closest in time.
ANITA: An optimal loopless accelerated variance-reduced gradient method
Li, Z · 2021
Closest in time.
ZeroSARAH: Efficient nonconvex finite-sum optimization with zero full gradient computation
Li, Z. and Richtárik, P · 2021
Closest in time.
EF21: A new, simpler, theoretically better, and practically faster error feedback
Richtárik, P., Sokolov, I., and Fatkhullin, I · 2021
Closest in time.