Fetching the paper…
Reading the bibliography…
Adaptivity is an important yet under-studied property in modern optimization theory.
Don’t jump through hoops and remove those loops: Svrg and katyusha are better without the outer loop
Kovalev, D., Horváth, S., and Richtárik, P · 1901
Earlier work this paper cites.
Stochastic newton and cubic newton methods with simple local linear-quadratic rates
Kovalev, D., Mishchenko, K., and Richtárik, P · 1912
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T · 1964
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovsky, A. S. and Yudin, D. B · 1983
Earlier work this paper cites.
Efficient estimators from a slowly convergent robbins-monro procedure
Ruppert, D · 1988
Earlier work this paper cites.
New stochastic approximation type procedures
Polyak, B. T · 1990
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B · 1992
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Moulines, E. and Bach, F. R · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Rakhlin, A., Shamir, O., and Sridharan, K · 2011
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence _rate for finite training sets
Roux, N. L., Schmidt, M., and Bach, F. R · 2012
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate o ( 1 / n ) o(1/n)
Bach, F. and Moulines, E · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Semi-stochastic gradient descent methods
Konečný, J. and Richtárik, P · 2013
Earlier work this paper cites.
Stochastic dual coordinate ascent methods for regularized loss minimization
Shalev-Shwartz, S. and Zhang, T · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Earlier work this paper cites.
A lower bound for the optimization of finite sums
Agarwal, A. and Bottou, L · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
From averaging to acceleration, there is only a step-size
Flammarion, N. and Bach, F · 2015
Earlier work this paper cites.
Variance reduced stochastic gradient descent with neighbors
Hofmann, T., Lucchi, A., Lacoste-Julien, S., and McWilliams, B · 2015
Cited alongside, same era.
Universal gradient methods for convex optimization problems
Nesterov, Y · 2015
Cited alongside, same era.
Quartz: Randomized dual coordinate ascent with arbitrary sampling
Qu, Z., Richtárik, P., and Zhang, T · 2015
Cited alongside, same era.
Stochastic optimization with importance sampling for regularized loss minimization
Zhao, P. and Zhang, T · 2015
Cited alongside, same era.
Less than a single pass: Stochastically controlled stochastic gradient method
Lei, L. and Jordan, M. I · 2016
Cited alongside, same era.
Stochastic variance reduction for nonconvex optimization
Reddi, S. J., Hefny, A., Sra, S., Poczos, B., and Smola, A · 2016
Lectures on convex optimization , volume 137
Nesterov, Y · 2018
Later among the works it cites.
k-svrg: Variance reduction for large scale optimization
Raj, A. and Stich, S. U · 2018
Later among the works it cites.
Local sgd converges fast and communicates little
Stich, S. U · 2018
Later among the works it cites.
Stochastic cubic regularization for fast nonconvex optimization
Tripuraneni, N., Stern, M., Jin, C., Regier, J., and Jordan, M. I · 2018
Later among the works it cites.
Spiderboost: A class of faster variance-reduced algorithms for nonconvex optimization
Wang, Z., Ji, K., Zhou, Y., Liang, Y., and Tarokh, V · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2017
Cited alongside, same era.
Katyusha: The first direct acceleration of stochastic gradient methods
Allen-Zhu, Z · 2017
Cited alongside, same era.
Harder, better, faster, stronger convergence rates for least-squares regression
Dieuleveut, A., Flammarion, N., and Bach, F · 2017
Cited alongside, same era.
S2CD: Semi-stochastic coordinate descent
Konečný, J., Qu, Z., and Richtárik, P · 2017
Cited alongside, same era.
Non-convex finite-sum optimization via scsg methods
Lei, L., Ju, C., Chen, J., and Jordan, M. I · 2017
Cited alongside, same era.
Distributed optimization with arbitrary local solvers
Ma, C., Konečný, J., Jaggi, M., Smith, V., Jordan, M. I., Richtárik, P., and Takáč, M · 2017
Cited alongside, same era.
Stochastic nested variance reduction for nonconvex optimization
Zhou, D., Xu, P., and Gu, Q · 2018
Later among the works it cites.
SGD: General analysis and improved rates
Gower, R. M., Loizou, N., Qian, X., Sailanbayev, A., Shulgin, E., and Richtárik, P · 2019
Later among the works it cites.
One method to rule them all: variance reduction for data, parameters and many new methods
Hanzely, F. and Richtárik, P · 2019
Later among the works it cites.
Revisiting the polyak step size
Hazan, E. and Kakade, S · 2019
Later among the works it cites.
A unified variance-reduced accelerated gradient method for convex optimization
Lan, G., Li, Z., and Zhou, Y · 2019
Later among the works it cites.
On the adaptivity of stochastic gradient-based optimization
Lei, L. and Jordan, M. I · 2019
Later among the works it cites.
Adaptive gradient descent without descent
Malitsky, Y. and Mishchenko, K · 2019
Later among the works it cites.
Finite-sum smooth optimization with sarah
Nguyen, L. M., van Dijk, M., Phan, D. T., Nguyen, P. H., Weng, T.-W., and Kalagnanam, J. R · 2019
Later among the works it cites.
On the convergence of adam and beyond
Reddi, S. J., Kale, S., and Kumar, S · 2019
Later among the works it cites.
Hybrid stochastic gradient descent algorithms for stochastic nonconvex optimization
Tran-Dinh, Q., Pham, N. H., Phan, D. T., and Nguyen, L. M · 2019
Later among the works it cites.
Painless stochastic gradient: Interpolation, line-search, and convergence rates
Vaswani, S., Mishkin, A., Laradji, I., Schmidt, M., Gidel, G., and Lacoste-Julien, S · 2019
Later among the works it cites.
Accelerate stochastic subgradient method by leveraging local growth condition
Xu, Y., Lin, Q., and Yang, T · 2019
Later among the works it cites.
Variance reduction with sparse gradients
Elibol, M., Lei, L., and Jordan, M. I · 2020
Closest in time.
A unified theory of sgd: Variance reduction, sampling, quantization and coordinate descent
Gorbunov, E., Hanzely, F., and Richtárik, P · 2020
Closest in time.