Fetching the paper…
Reading the bibliography…
Techniques for reducing the variance of gradient estimates used in stochastic programming algorithms for convex finite-sum problems have received a great deal of attention in recent years.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
The heavy ball with friction method, i. the continuous dynamical system: global exploration of the local minima of a real-valued function by asymptotic analysis of a dissipative dynamical system
Attouch, H., Goudou, X., and Redont, P · 2000
Earlier work this paper cites.
A second-order gradient-like dissipative dynamical system with hessian-driven damping.-application to optimization and mechanics
Alvarez, F., Attouch, H., Bolte, J., and Redont, P · 2002
Earlier work this paper cites.
Large scale online learning
Bottou, L. and LeCun, Y · 2003
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Nesterov, Y · 2003
Earlier work this paper cites.
Dissipative dynamical systems
Willems, J · 2007
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for strongly-convex optimization with finite training sets
Roux, N., Schmidt, M., and Bach, F · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Minimizing finite sums with the stochastic average gradient
Schmidt, M., Roux, N., and Bach, F · 2013
Earlier work this paper cites.
Stochastic dual coordinate ascent methods for regularized loss
Shalev-Shwartz, S. and Zhang, T · 2013
Earlier work this paper cites.
Linear coupling: An ultimate unification of gradient and mirror descent
Allen-Zhu, Z. and Orecchia, L · 2014
Earlier work this paper cites.
An accelerated proximal coordinate gradient method
Lin, Q., Lu, Z., and Xiao, L · 2014
Cited alongside, same era.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
Needell, D., Ward, R., and Srebro, N · 2014
Cited alongside, same era.
Convex optimization: Algorithms and complexity
Bubeck, S · 2015
Cited alongside, same era.
A geometric alternative to Nesterov’s accelerated gradient descent
Bubeck, S., Lee, Y., and Singh, M · 2015
Cited alongside, same era.
A universal catalyst for first-order optimization
Lin, H., Mairal, J., and Harchaoui, Z · 2015
Cited alongside, same era.
Katyusha: The first direct acceleration of stochastic gradient methods
SDCA without duality, regularization, and individual convexity
Shalev-Shwartz, S · 2016
Later among the works it cites.
A differential equation for modeling nesterov’s accelerated gradient method: Theory and insights
Su, W., Boyd, S., and Candès, E · 2016
Later among the works it cites.
Barzilai-borwein step size for stochastic gradient descent
Tan, C., Ma, S., Dai, Y., and Qian, Y · 2016
Later among the works it cites.
A variational perspective on accelerated methods in optimization
Wibisono, A., Wilson, A., and Jordan, M · 2016
Later among the works it cites.
A lyapunov analysis of momentum methods in optimization
Wilson, A., Recht, B., and Jordan, M · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Allen-Zhu, Z · 2016
Cited alongside, same era.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F., and Nocedal, J · 2016
Cited alongside, same era.
A simple practical accelerated method for finite sums
Defazio, A · 2016
Cited alongside, same era.
An optimal first order method based on optimal quadratic averaging
Drusvyatskiy, D., Fazel, M., and Roy, S · 2016
Cited alongside, same era.
Analysis and design of optimization algorithms via integral quadratic constraints
Lessard, L., Recht, B., and Packard, A · 2016
Cited alongside, same era.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S
Cited in the paper.
Finito: A faster, permutable incremental gradient method for big data problems
Defazio, A., Domke, J., and Caetano, T
Cited in the paper.
Fazlyab, M., Ribeiro, A., Morari, M., and Preciado, V · 2017
Later among the works it cites.
Dissipativity theory for Nesterov’s accelerated method
Hu, B. and Lessard, L · 2017
Later among the works it cites.
SARAH: A novel method for machine learning problems using stochastic recursive gradient
Nguyen, L., Liu, J., Scheinberg, K., and Takáč, M · 2017
Later among the works it cites.
Stochastic primal-dual coordinate method for regularized empirical risk minimization
Zhang, Y. and Xiao, L · 2017
Later among the works it cites.
Katyusha x: Practical momentum method for stochastic sum-of-nonconvex optimization
Allen-Zhu, Z · 2018
Closest in time.